Correlithm Object Processing for Data Clustering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computers are limited in comparing and determining similarity between data samples, relying on complex signal processing techniques due to the ordinal nature of their number systems, which consumes processing power and reduces performance, especially in applications like face recognition and fraud detection.
Innovation Solution
The implementation of a correlithm object processing system that uses categorical numbers and geometric objects to represent data samples, enabling non-binary comparisons and quantifying similarity between data samples, regardless of their type or format, through the use of correlithm objects and a combination of sensor, node, and actor tables.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional computers use ordinal binary integers to represent and manipulate data samples, then they can perform basic operations such as counting, sorting, and indexing, but they cannot efficiently determine similarity between different data samples without complex signal processing techniques
Solution Approach 1:
The patent transforms the number system parameter from ordinal binary integers to a hybrid system incorporating categorical numbers and geometric correlithm objects. This parameter change enables direct similarity measurement through geometric operations (comparing positions and orientations in multidimensional space) rather than complex signal processing, thereby improving similarity detection capability while reducing processing complexity
Solution Approach 2:
The patent introduces correlithm objects as intermediary geometric representations between raw data samples and similarity comparisons. These correlithm objects serve as mediators that encode data characteristics in geometric form, allowing the system to determine similarity through geometric operations rather than direct complex signal processing between original data samples
2Measurement precision
If conventional computers rely on complex signal processing techniques to compare data samples, then they can determine similarity, but processing power is consumed and system performance is reduced
Solution Approach 1:
The patent replaces the mechanical/computational signal processing system with a geometric processing system. Instead of using complex algorithms and computational operations to compare data samples, the system uses geometric operations (comparing positions, orientations, and distances of correlithm objects in multidimensional space), which are computationally simpler and faster while maintaining similarity detection accuracy
3Adaptability or versatility
If conventional computers use ordinal numbers to represent data samples, then they can store and manipulate information, but they are unable to tell if a data sample matches or is similar to other data samples unless there is an exact match
Solution Approach 1:
The patent transitions from one-dimensional ordinal number representation to multidimensional geometric representation using correlithm objects. Each correlithm object exists in multidimensional space with specific position and orientation characteristics that encode data sample features. This dimensional expansion allows the system to capture and represent similarity information that cannot be expressed in ordinal number systems, enabling flexible matching and similarity detection
Data Source
AI summary
A device that includes a model training engine implemented by a processor. The model training engine is configured to obtain a set of data values associated with a feature vector. The model training engine is further configured to generate a set of gradients by dividing separation distances by an average separation distance and to compare each gradient to a gradient threshold value. The model training engine is further configured to identify a boundary in response to determining a gradient exceeds the gradient threshold value, to determine a number of identified boundaries, and to determine a number of clusters based on the number of identified boundaries. The model training engine is further configured to train the machine learning model to associate the determined number of clusters with the feature vector.


