Unstructured Data Search via Subspace Projection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for searching and displaying unstructured data face challenges such as polysemy and synonymy, requiring significant computational resources and inefficient results, especially when dealing with large collections of documents like network flow data.
Innovation Solution
A system that utilizes subspace representation by retrieving tokens, associating them with multidimensional coordinates, and creating a normalized matrix to enable efficient search and display of unstructured data, allowing for timely detection and mitigation of cyberattacks through the use of attack-specific markers in network traffic data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional vector-space models are used for information retrieval, then search coverage is achieved, but polysemy and synonymy problems occur leading to retrieval of unrelated documents or missing related documents
Solution Approach 1:
The patent transforms the traditional vector-space model by introducing a subspace dimension. Instead of working in the full high-dimensional space where polysemy and synonymy cause problems, the system projects data into a lower-dimensional subspace that captures the essential semantic relationships. This dimensional transformation allows the system to maintain search coverage while improving retrieval precision by eliminating the harmful effects of polysemy and synonymy that plague traditional models.
Solution Approach 2:
The system changes the parameters of the search space by transforming the original high-dimensional vector space into a reduced subspace with different dimensional characteristics. This parameter change involves selecting a subset of dimensions that are most relevant to the search query, thereby altering the search parameters to achieve both comprehensive coverage and high precision in retrieving related documents.
2Productivity
If subspace representation is used to improve search efficiency, then computational resources are reduced, but the system must handle very large bodies of document data and tokens
Solution Approach 1:
The patent extracts only the essential and relevant features from very large bodies of document data and tokens by projecting them into a subspace. Instead of processing the complete high-dimensional data set, the system extracts the most important dimensional components that capture the semantic essence of the data. This extraction process reduces computational resources required while maintaining the ability to handle large volumes of data effectively.
Solution Approach 2:
By transforming the data into a lower-dimensional subspace, the system reduces the quantity of computational operations needed while preserving the essential information. The dimensional reduction allows the system to process large volumes of documents and tokens efficiently by working in a compressed representation that retains the critical semantic relationships needed for accurate search.
3Loss of energy
If modest storage and computational resources are used, then resource efficiency is improved, but the system must still organize and search massive collections of electronic documents
Solution Approach 1:
The patent creates a compressed subspace copy or projection of the massive document collection that preserves the essential searchability and organizational structure. Instead of storing and processing the complete high-dimensional data, the system maintains a reduced-dimensional representation that acts as an efficient copy for search and organization operations. This copying approach enables modest resources to handle massive collections by working with the essential structural information rather than the full data complexity.
Solution Approach 2:
The system changes the storage and computational parameters by representing documents in a reduced subspace with fewer dimensions. This parameter change reduces the storage requirements and computational overhead while maintaining the capability to organize and search massive document collections effectively. The transformed parameters enable resource-efficient processing without sacrificing organizational capability.
Data Source
AI summary
A system to organize, search and display unstructured data comprising a token retrieval module, a document indexing engine, a subspace search module and a user interface module has been devised. The system retrieves a plurality of tokens and associates them with coordinates in subspace. It also retrieves documents and creates a multidimensional matrix of documents and tokens where each cell contains the number of times the token occurs in each document. That matrix is employed in a search using user specified search terms. The search results are displayed such that the search tokens occupy specific spatial coordinates and documents spatial coordinates are dictated by the relative preponderance of each search term in each document.


