User Clustering via Centroid Vectors for Privacy Compliance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in accurately clustering users across different environments while respecting privacy constraints imposed by varying privacy policies and regulations, where characteristics information from one environment cannot be directly obtained or shared with another.
Innovation Solution
A processing system comprising a first and second processing device, where the first device divides users into clusters based on first characteristics information and transmits centroid vector representation to the second device, which then clusters users using the received representation information and its own second characteristics information, ensuring privacy compliance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If characteristics information from multiple environments is combined to improve clustering accuracy, then clustering accuracy is improved, but privacy compliance deteriorates due to restrictions imposed by different privacy policies and regulations
Solution Approach 1:
The patent segments the characteristics information processing into separate environments. Each environment (first environment and second environment) processes its own characteristics information independently to produce cluster labels. This segmentation allows clustering accuracy to be improved by considering multiple environments while maintaining privacy compliance, as each environment processes data according to its own privacy policies without sharing raw characteristics information.
Solution Approach 2:
The patent introduces cluster labels as an intermediary between different environments. Instead of directly sharing raw characteristics information, each environment uses its characteristics information to generate cluster labels, which then serve as the basis for combining results across environments. This intermediary approach enables accurate user clustering while respecting privacy restrictions in each environment.
2Measurement precision
If characteristics information is shared between different environments to enable comprehensive user analysis, then analysis accuracy is improved, but data security deteriorates due to privacy policies and regulations
Solution Approach 1:
The patent extracts only the necessary cluster labels from each environment rather than sharing the complete characteristics information. Each environment extracts and transmits only the cluster labels derived from its own characteristics information, eliminating the need to share sensitive raw data while still enabling comprehensive user analysis across environments.
Solution Approach 2:
The patent creates a simplified representation (copy) of user characteristics through cluster labels. Instead of sharing the original characteristics information, each environment generates a simplified cluster label representation that captures essential user characteristics for analysis purposes, reducing data security risks while maintaining analysis accuracy.
3Productivity
If raw characteristics information is exchanged between processing devices, then clustering can be performed using all available data, but privacy protection deteriorates
Solution Approach 1:
The patent introduces cluster labels as an intermediary that enables clustering capability across multiple environments without exchanging raw characteristics information. Each processing device generates cluster labels from its local characteristics information, and these labels serve as the basis for coordinated clustering while preserving privacy protection by preventing raw data exchange.
Solution Approach 2:
The patent extracts only the essential clustering information (cluster labels) from each environment and uses this extracted information for coordinated clustering. This extraction approach maintains productivity by enabling multi-environment clustering while protecting privacy by preventing the exchange of raw characteristics information.
Data Source
AI summary
In a processing system 101, a first processing device 111 obtains first characteristics information indicating characteristics of users in a first environment, divides the users into first clusters according to the first characteristics information, and transmits, to a second processing device 121, representation information representing characteristics of users belonging to each of the first clusters using a centroid vector based on the first characteristics information of the cluster. The second processing device 121 obtains second characteristics information indicating characteristics of the users in a second environment, and divides the users into second clusters according to the representation information and the second characteristics information. The present invention allows for dividing the users into clusters more accurately, although the second processing device 121 need not use the first characteristics information and the first processing device 111 need not use the second characteristics information.


