Self-Organizing Map for Unstructured Data Clustering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data search and analysis tools are inadequate for efficiently organizing and interpreting large amounts of structured and unstructured data, particularly in contexts like intelligence gathering, where relevant information may be overlooked due to the difficulty in accessing and understanding unstructured data from diverse sources.

Innovation Solution

A computer-implemented method and system that converts textual data into numeric representations, forms self-organizing maps to cluster similar data, and generates dialectic arguments to interpret and synthesize the data, facilitating the organization and analysis of vast information sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual scanning of available facts and figures is used to search for useful information, then comprehensive review of data is possible, but the process takes considerable time and effort and often overlooks key elements

Engineering Contradiction:
Improvecompleteness of information retrievalVSAvoidtime and effort for data analysis
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical scanning of data with automated computational systems including search engines, data mining tools, and neural networks that can process vast amounts of structured and unstructured data automatically, dramatically reducing time and effort while maintaining or improving completeness of information retrieval

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates multiple copies and representations of data including full-text indexes, structured extracts, and visual concept maps that allow simultaneous analysis from different perspectives, enabling comprehensive review without repeated manual scanning of original sources

Inventive Principle:
Principle #26Copying

2Productivity

If search engines are used to electronically interrogate available data sources, then large amounts of data can be accessed quickly, but the unstructured data remains difficult to organize, search, and retrieve useful information

Engineering Contradiction:
Improvespeed of data accessVSAvoiddifficulty of organizing and retrieving unstructured data
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent introduces intermediate processing layers including full-text indexing, natural language processing, and automated concept extraction systems that mediate between raw unstructured data and user queries, transforming difficult-to-search text into organized, queryable formats while preserving the original data

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the parameters of unstructured data by converting text into multiple representations including keyword indexes, structured metadata, and visual concept maps, making the data searchable and retrievable while maintaining the original unstructured content for reference

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If multiple government computer systems are used for intelligence gathering, then vast amounts of structured information are available, but the systems are generally not linked together and data from one agency is not necessarily available to another

Engineering Contradiction:
Improveamount of available informationVSAvoidcomplexity of data integration
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements universal data access protocols and standardized interfaces that allow different agency systems to share and access each other's data through common mechanisms, enabling multi-functional use of data across different intelligence gathering operations without requiring separate systems for each agency

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent creates a nested architecture where individual agency databases are embedded within a larger federated system, allowing data from one agency to be accessed by another while maintaining the organizational structure and security boundaries of each individual system

Inventive Principle:
Principle #7Nested doll (Nesting)

4Ease of operation

If unstructured data is stored in document format, then the data can be easily stored and accessed by those in possession of the document, but the data has little meaning beyond immediate context and is difficult to interpret

Engineering Contradiction:
Improveease of storage and accessVSAvoidloss of meaning and context
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent adds another dimension to unstructured data by creating visual concept maps and structured representations that display relationships and meanings beyond the immediate textual context, allowing users to interpret data meaning while preserving the original document format for storage and access

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS7447665B2System and method of self-learning conceptual mapping to organize and interpret data
Publication Date: 2008.11.04 KINETX
  • US7447665B2 patent drawing
  • US7447665B2 patent drawing
  • US7447665B2 patent drawing

AI summary

In a computer implemented method of researching textual data sources, textual data is reduced to a plurality of distinctive words based on frequency of usage within the textual data. The distinctive words are converted into first numeric representations of vectors containing random numbers. A first self-organizing map is formed from the first numeric representations and organized by similarities between the vectors. A second self-organizing map is formed from second numeric representations generated from the organization of the first self-organizing map. The second numeric representations are vectors derived from the first self-organizing map. The vectors are used to train the second self-organizing map. The vectors derived from the first self-organizing map are organized into clusters of similarities between the vectors on the second self-organizing map. Dialectic arguments are formed from the second self-organizing map to interpret the textual data.