Cross-Environment User Clustering via Cluster Label Mediation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge lies in accurately clustering users across different environments while respecting privacy constraints imposed by varying privacy policies, agreements, or regulations, where characteristics information from one environment cannot be directly obtained or used in another.
Innovation Solution
A processing system comprising a first and second processing device, where the first device divides users into clusters based on first characteristics information and transmits belonging information, which the second device uses to divide users into separate clusters using second characteristics information, ensuring privacy is maintained by not exchanging raw data between environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If characteristics information from multiple environments is directly combined for clustering, then clustering accuracy is improved, but privacy compliance deteriorates
Solution Approach 1:
The patent introduces an intermediary mechanism where cluster labels from the first environment are transmitted to the second environment, serving as a bridge that enables cross-environment clustering without directly sharing sensitive characteristics information. This mediator (cluster labels) allows the system to achieve accurate clustering while maintaining privacy compliance, as it transmits only aggregated group identifiers rather than raw user data.
Solution Approach 2:
The patent segments the clustering process into two independent parts: first environment clustering that produces cluster labels, and second environment clustering that uses these labels as references. This segmentation allows each environment to perform clustering separately on its own data while still achieving coordinated results, thereby maintaining privacy while improving accuracy.
2Object-affected harmful factors
If characteristics information is kept separate in different environments, then privacy compliance is improved, but clustering accuracy deteriorates
Solution Approach 1:
The patent implements a feedback mechanism where cluster labels generated in the first environment are fed into the second environment's clustering process as reference information. This feedback loop enables the second environment to adjust its clustering to be consistent with the first environment's results, thereby maintaining privacy compliance while improving overall clustering accuracy through iterative refinement.
3Stability of the object's composition
If raw characteristics information is exchanged between processing devices, then clustering coordination is improved, but data security deteriorates
Solution Approach 1:
The patent uses copying by transmitting only cluster labels (aggregated group identifiers) between processing devices rather than raw characteristics information. These labels are simplified copies that represent clustering results without containing sensitive user data, thereby maintaining clustering consistency while preserving data security through information abstraction.
Data Source
AI summary
In a processing system 101, a first processing device 111 obtains first characteristics information indicating a characteristics of users in a first environment, divides the users into first clusters according to the first characteristics information, and transmits, to a second processing device 121, belonging information indicating which of the first clusters each of the users belongs to. The second processing device 121 obtains second characteristics information indicating characteristics of the users in a second environment, and divides the users into second clusters according to the belonging information and the second characteristics information. The present invention allows for dividing the users into clusters more accurately while keeping the first and second characteristics information separate, although the second processing device 121 need not use the first characteristics information collected in the first environment and the first processing device 111 need not use the second characteristics information collected in the second environment.


