Mixture Model Clustering for Data Without Feature Information

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing clustering methods, such as the general mixture model and multivariate regression trees, are inefficient in classifying new stores without sales information and struggle with data represented by continuous values, limiting their ability to generate document clusters effectively.

Innovation Solution

A clustering system using a mixture model defined by two types of variables, where the mixing ratio is represented by a function of a first variable and the element distribution of the cluster is represented by a function of a second variable, allowing for classification of target data independently of whether it has information indicating the features of a cluster.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a general mixture model is used for clustering, then existing data with feature information can be classified, but new stores without sales information cannot be appropriately clustered

Engineering Contradiction:
Improveclustering capabilityVSAvoidclassification accuracy for new data
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the mixture model into two distinct functional components: mixing ratio functions that operate without feature information and element distribution functions that utilize feature information when available. This segmentation allows the system to handle both new stores (using only mixing ratio) and existing stores (using both components) effectively.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal clustering framework where the mixture model can operate in multiple modes: with feature information (using both mixing ratio and element distribution functions) and without feature information (using only mixing ratio function). This multi-functionality resolves the contradiction by making the system adaptable to different data availability scenarios while maintaining reliable classification.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multivariate regression trees are used for data division, then continuous value data can be processed, but data represented by non-continuous values (e.g., document clusters) cannot be effectively generated

Engineering Contradiction:
Improvedata type handlingVSAvoidclustering effectiveness
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent changes the mathematical parameters of the clustering model from regression-based continuous value processing to probability-based discrete cluster assignment. By using mixing ratio functions that output probabilities and element distribution functions that model discrete data generations, the system can effectively handle non-continuous data types like document clusters while maintaining high clustering effectiveness.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If clustering is performed based on sales feature vectors, then stores with sales information can be segmented, but new stores without sales information cannot be classified

Engineering Contradiction:
Improveclustering precision for existing storesVSAvoidapplicability to new stores
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary action by pre-defining mixing ratio functions that can operate independently of feature information. These functions are prepared in advance to handle cases where feature data is unavailable, allowing new stores to be classified immediately upon arrival without requiring sales information, while existing stores still benefit from precise feature-based clustering.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10877996B2Clustering system, method, and program
Publication Date: 2020.12.29 NEC CORP
  • US10877996B2 patent drawing
  • US10877996B2 patent drawing
  • US10877996B2 patent drawing

AI summary

A classifier 81 classifies target data into a cluster on the basis of a mixture model defined using two different types of variables that indicate features of the target data. In this classification, the classifier 81 classifies the target data into a cluster on the basis of a mixture model in which a mixing ratio of the mixture model is represented by a function of a first variable and in which the element distribution of the clusters into which the target data is classified is represented by a function of a second variable.