Mixture Model Clustering for Data Without Feature Information
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing clustering methods, such as the general mixture model and multivariate regression trees, are inefficient in classifying new stores without sales information and struggle with data represented by continuous values, limiting their ability to generate document clusters effectively.
Innovation Solution
A clustering system using a mixture model defined by two types of variables, where the mixing ratio is represented by a function of a first variable and the element distribution of the cluster is represented by a function of a second variable, allowing for classification of target data independently of whether it has information indicating the features of a cluster.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a general mixture model is used for clustering, then existing data with feature information can be classified, but new stores without sales information cannot be appropriately clustered
Solution Approach 1:
The patent segments the mixture model into two distinct functional components: mixing ratio functions that operate without feature information and element distribution functions that utilize feature information when available. This segmentation allows the system to handle both new stores (using only mixing ratio) and existing stores (using both components) effectively.
Solution Approach 2:
The patent creates a universal clustering framework where the mixture model can operate in multiple modes: with feature information (using both mixing ratio and element distribution functions) and without feature information (using only mixing ratio function). This multi-functionality resolves the contradiction by making the system adaptable to different data availability scenarios while maintaining reliable classification.
2Adaptability or versatility
If multivariate regression trees are used for data division, then continuous value data can be processed, but data represented by non-continuous values (e.g., document clusters) cannot be effectively generated
Solution Approach 1:
The patent changes the mathematical parameters of the clustering model from regression-based continuous value processing to probability-based discrete cluster assignment. By using mixing ratio functions that output probabilities and element distribution functions that model discrete data generations, the system can effectively handle non-continuous data types like document clusters while maintaining high clustering effectiveness.
3Measurement precision
If clustering is performed based on sales feature vectors, then stores with sales information can be segmented, but new stores without sales information cannot be classified
Solution Approach 1:
The patent performs preliminary action by pre-defining mixing ratio functions that can operate independently of feature information. These functions are prepared in advance to handle cases where feature data is unavailable, allowing new stores to be classified immediately upon arrival without requiring sales information, while existing stores still benefit from precise feature-based clustering.
Data Source
AI summary
A classifier 81 classifies target data into a cluster on the basis of a mixture model defined using two different types of variables that indicate features of the target data. In this classification, the classifier 81 classifies the target data into a cluster on the basis of a mixture model in which a mixing ratio of the mixture model is represented by a function of a first variable and in which the element distribution of the clusters into which the target data is classified is represented by a function of a second variable.


