Confidential Data Backend for Smoothed Posterior Distribution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

There is a challenge in collecting and maintaining confidential data in computer systems while ensuring user privacy, as users are hesitant to share sensitive information due to concerns about data security and misuse, particularly in balancing the quality of statistical insights with data coverage.

Innovation Solution

A system is developed that securely collects and processes confidential data by using a confidential data frontend and backend architecture, employing encryption techniques to separate user identification from confidential data, anonymizing data points, and employing log-linear models to estimate posterior distributions for cohorts with small sample sizes, thereby enhancing data privacy and quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If confidential data is collected and stored in a computer system, then statistical insights and data coverage are improved, but user privacy and data security are compromised

Engineering Contradiction:
Improvestatistical insightsVSAvoidprivacy concerns
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The patent introduces a confidential data backend as an intermediary component that acts as a trusted mediator between users and the data collection system. This backend securely stores confidential data, performs statistical computations, and returns only aggregated insights without exposing raw data, thus resolving the contradiction between gaining statistical value and protecting user privacy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent extracts and separates user identification information from confidential data through anonymization techniques. By removing or masking identifying elements, the system maintains statistical utility while eliminating privacy risks, allowing data to be used for insights without compromising user identity

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If groupings are selected from a larger number of samples, then the quality of statistical insights is improved, but the coverage is reduced

Engineering Contradiction:
Improvequality of statistical insightsVSAvoiddata coverage
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent implements dynamic grouping strategies that can adjust the granularity and composition of data groupings based on available sample sizes and statistical requirements. This allows the system to optimize between insight quality and coverage by adapting groupings to current data conditions rather than using fixed groupings

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent merges multiple data groupings and samples through the confidential data backend, combining information from various cohorts and segments. This merging approach enables the system to maintain comprehensive coverage while still producing high-quality statistical insights by aggregating data across multiple groups

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS10552741B1Computing smoothed posterior distribution of confidential data
Publication Date: 2020.02.04 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10552741B1 patent drawing
  • US10552741B1 patent drawing
  • US10552741B1 patent drawing

AI summary

In an example, a set of cohort types and an anonymized set of confidential data data values for a plurality of cohorts having cohort types in the set of cohort types are obtained. Then it is determined, from a set of candidate data transformations, a best fitting data transformation for the anonmyized set of confidential data data values. The anonymized set of confidential data data values is transformed using the best fitting data transformation. Optimal smoothing parameters are computed for each cohort type. Then, for each cohort in the set of cohort types having a small sample size, a best parent for the cohort is determined and a posterior distribution for the cohort is determined based on the best parent for the cohort and the optimal smoothing parameters for a cohort type for the cohort.