Correlithm Object Clustering for Data Similarity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computers are limited in comparing and determining similarity between data samples, relying on complex signal processing techniques due to the ordinal nature of their number systems, which consumes processing power and reduces performance, especially in applications like face recognition and fraud detection.
Innovation Solution
The implementation of a correlithm object processing system that uses categorical numbers and geometric objects to represent data samples, enabling non-binary comparisons and quantifying similarity between data samples, regardless of their type or format, through the use of sensor, node, and actor tables.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional computers use ordinal binary integers to represent and manipulate data samples, then they can perform operations such as counting, sorting, indexing, and mathematical calculations, but they cannot determine similarity between different data samples without complex signal processing techniques
Solution Approach 1:
The patent transforms the numerical representation system from ordinal binary integers to a correlation-based number system. Data samples are represented by correlithm objects with correlithm values that encode similarity relationships directly. This parameter change in the number system enables direct similarity determination without complex signal processing, as the correlithm values inherently contain similarity information through their mathematical properties.
2Measurement precision
If conventional computers rely on complex signal processing techniques to determine similarity between data samples, then they can achieve accurate comparison, but processing power is consumed and system performance is reduced
Solution Approach 1:
The patent extracts the similarity determination function from complex signal processing operations and embeds it directly into the data representation structure. By representing data samples as correlithm objects with correlithm values that mathematically encode similarity relationships, the system extracts the essential similarity information from the data itself, enabling fast comparison operations without requiring separate complex processing stages.
Solution Approach 2:
The patent performs preliminary transformation of data samples into correlithm object representations during data ingestion or preprocessing. This preliminary action encodes similarity relationships into the correlithm values before comparison operations are needed, so that subsequent similarity determinations can be performed rapidly by simply comparing correlithm values rather than executing complex signal processing algorithms in real-time.
3Adaptability or versatility
If conventional computers use ordinal numbers to represent data samples, then they can store and manipulate information efficiently, but they are unable to tell if a data sample matches or is similar to any other data samples unless there is an exact match
Solution Approach 1:
The patent fundamentally changes the parameter system from ordinal numbers to correlithm values. Ordinal numbers only provide sequence information, while correlithm values are specifically designed to encode similarity relationships. This parameter change enables the system to identify both exact matches and similar data samples through direct correlithm value comparison, making the comparison operation simpler and more versatile.
Data Source
AI summary
A device comprising a cluster engine implemented by a processor. The cluster engine is configured to obtain a reference correlithm object and compute a set of Anti-Hamming distances between the reference correlithm object and the set of correlithm objects. The cluster engine is further configured to identify a subset of correlithm objects from the set of correlithm objects that are associated with an Anti-Hamming distance that is greater than a first bit threshold value. The cluster engine is further configured to compute a set of Hamming distances between the reference correlithm object and the subset of correlithm objects and to identify correlithm objects associated with a Hamming distance that exceeds a second bit threshold value. The cluster engine is further configured to remove the identified correlithm objects that are associated with a Hamming distance that exceeds the second bit threshold value and generate the cluster.


