Unstructured Data Pattern Detection via Signature Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for analyzing unstructured data from big data sources are inefficient in identifying common patterns due to the complexity and lack of correlation between data elements, leading to inefficient search and analysis processes.

Innovation Solution

A method and system for extracting unstructured data elements, generating robust signatures, clustering these signatures to identify common patterns, and correlating clusters to detect associations between patterns, utilizing a network interface, processor, and memory to facilitate efficient analysis and correlation within big data sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If traditional data processing applications are used to analyze unstructured data from big data sources, then the analysis can be performed with simple tools, but the identification of common patterns becomes inefficient due to data complexity

Engineering Contradiction:
Improvesimplicity of analysis toolsVSAvoidefficiency of pattern identification
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent segments unstructured data into structured formats by extracting specific data elements and organizing them into standardized schemas. This segmentation transforms complex unstructured data into manageable structured components that can be efficiently analyzed for common patterns while maintaining tool simplicity.

Inventive Principle:
Principle #1Segmentation

2Ease of operation

If data elements are extracted from big data sources without correlation, then the extraction process is straightforward, but the search for additional useful data becomes inefficient

Engineering Contradiction:
Improveease of data extractionVSAvoidtime for searching additional data
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent implements feedback mechanisms where extracted data elements are correlated with previously extracted elements to identify relationships and patterns. This feedback loop enables the system to learn from extracted data and improve subsequent extraction efficiency, reducing time for searching additional useful data while maintaining ease of extraction.

Inventive Principle:
Principle #23Feedback

3Device complexity

If unstructured data is analyzed without signature generation, then the analysis process is simpler, but the identification of common patterns among data elements becomes inefficient

Engineering Contradiction:
Improvecomplexity of analysis processVSAvoidefficiency of common pattern identification
Core Design Contradiction:
Device complexityVSProductivity

Solution Approach 1:

The patent generates signatures that are simplified representations or copies of complex unstructured data elements. These signatures capture essential characteristics of the original data in a condensed format, enabling efficient pattern identification without requiring analysis of the full complexity of the original unstructured data.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9256668B2System and method of detecting common patterns within unstructured data elements retrieved from big data sources
Publication Date: 2016.02.09 CORTICA LTD
  • US9256668B2 patent drawing
  • US9256668B2 patent drawing
  • US9256668B2 patent drawing

AI summary

A method for detection of common patterns within unstructured data elements. The method includes extracting a plurality of unstructured data elements retrieved from a plurality of big data sources; generating at least one signature for each of the plurality of unstructured data elements; identifying common patterns among the generated signatures; clustering the signatures identified to have common patterns; and correlating the generated clusters to identify associations between their respective identified common patterns.