Learning Device Latent Vector Clustering Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional learning technologies face challenges in accurately clustering and classifying complex data, such as image data with diverse backgrounds, due to difficulties in calculating distance and similarity, leading to decreased clustering accuracy.
Innovation Solution
A learning device that calculates latent vectors using a machine learning model, updates parameters based on both first and second loss functions to enhance clustering accuracy, and classifies data by unsupervised clustering methods, improving the separation of data points in the latent space.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Volume of moving object
If conventional learning methods are used to represent complex data by low-dimensional feature vectors, then dimensionality reduction is achieved, but clustering accuracy decreases due to difficulties in calculating distance and similarity
Solution Approach 1:
The patent changes the parameter representation from conventional Euclidean distance to cosine similarity measurement in the latent space. By transforming the distance calculation metric and introducing virtual class probabilities, the system maintains clustering accuracy while working with low-dimensional feature vectors. The loss function is also modified to incorporate both clustering objectives and virtual class classification, changing the optimization parameters to achieve better results.
2Productivity
If data with diverse backgrounds (e.g., image data) are clustered using conventional methods, then processing speed is maintained, but classification accuracy decreases
Solution Approach 1:
The patent introduces virtual classes as an intermediary concept between the input data and final clustering results. By assuming data belong to virtual classes and calculating probabilities of belonging to these virtual classes, the system creates an intermediate representation that guides the clustering process. This intermediary mechanism enables accurate classification of diverse data while maintaining processing efficiency through the learned latent space representation.
3Measurement precision
If supervised data are used for learning, then classification accuracy improves, but the requirement for labeled data increases learning complexity
Solution Approach 1:
The patent implements self-service learning by formulating the clustering objective as a self-supervised learning problem. The system automatically creates virtual class labels from the data itself and uses these to train the model without requiring external supervised data. The loss function combines clustering objectives with virtual class classification, enabling the model to learn useful representations autonomously from unlabeled data, thus reducing learning complexity while maintaining accuracy.
Data Source
AI summary
According to an embodiment, a learning device includes one or more processors. The processors calculate a latent vector of each of a plurality of first target data, by using a parameter of a learning model configured to output a latent vector indicating a feature of a target data. The processors calculate, for each first target data, first probabilities that the first target data belongs to virtual classes on an assumption that the plurality of first target data belong to the virtual classes different from each other. The processors update the parameter such that a first loss of the first probabilities, and a second loss that is lower as, for each of element classes to which a plurality of elements included in each of the plurality of first target data belong, a relation with another element class is lower, become lower.


