Sparse Distributed Representations for Cross-Lingual Semantic Mapping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems for data document clustering and semantic analysis do not effectively utilize self-organizing maps to generate cross-lingual sparse distributed representations (SDRs) for explicit semantic definition of data items.
Innovation Solution
A method that involves clustering data documents in a two-dimensional metric space using a reference map generator, generating semantic maps, and creating sparse distributed representations (SDRs) for terms based on their occurrence information across documents, allowing for the identification of similarity between data items and filtering criteria.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional self-organizing maps are used for data document clustering, then data documents can be clustered by type, but the system cannot generate cross-lingual sparse distributed representations for explicit semantic definition of data items
Solution Approach 1:
The self-organizing map is extended to perform multiple functions: traditional data document clustering by type, and generation of sparse distributed representations for explicit semantic definition. This multi-functional approach enables cross-lingual semantic representation while utilizing the existing clustering infrastructure.
Solution Approach 2:
The system segments the semantic representation task into discrete sparse distributed representations for individual data items, allowing each item to have its own explicit semantic definition while maintaining overall system coherence through the self-organizing map structure.
2Measurement precision
If sparse distributed representations are generated for explicit semantic definition, then semantic analysis capability is improved, but computational complexity increases
Solution Approach 1:
The system changes the parameter representation from traditional vector spaces to sparse distributed representations, which encode semantic information in a distributed manner across multiple units. This transformation enables precise semantic similarity measurement through comparison of SDR patterns while the sparsity constraint manages computational complexity.
Data Source
AI summary
A method enables identification of a similarity level between a user-provided data item and a data item within a set of data documents. The method includes a representation generator determining, for each term in an enumeration of terms, occurrence information. The representation generator generates, for each term, a sparse distributed representation (SDR) using the occurrence information. The method includes receiving, by a filtering module, a filtering criterion. The method includes generating, by the representation generator, for the filtering criterion, at least one SDR. The method includes generating, by the representation generator, for a first of a plurality of streamed documents received from a data source, a compound SDR. The method includes determining, by a similarity engine executing on the second computing device, a distance between the filtering criterion SDR and the generated compound SDR. The method includes acting on the first streamed document, based upon the determined distance.


