Infocube Data Shuffling for Secure Testing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In the context of SAP's Business Warehouse system, using actual production data for testing poses a risk of sensitive information disclosure, and relying on artificial data limits the effectiveness of testing, necessitating a mechanism to create anonymous data for secure testing purposes.

Innovation Solution

A method and apparatus for shuffling data within infocubes, specifically rearranging columns of fact, dimension, and characteristic tables using random number assignment or dictionaries, to make the data anonymous and protect sensitive information, enabling its use in testing applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If actual production data is used for testing, then testing effectiveness is improved, but risk of sensitive information disclosure increases

Engineering Contradiction:
Improvetesting effectivenessVSAvoidsensitive information disclosure risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates a copy of the actual production data through infocubes, then applies shuffling operations to this copy. The shuffled infocube data maintains the statistical properties and relationships of the original data for effective testing, while the values are permuted to prevent identification of sensitive information. This resolves the contradiction by using a transformed copy rather than the original data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the data by changing the parameter arrangement through shuffling operations on infocube values. By permuting the values while maintaining their distribution characteristics, the data remains useful for testing statistical functions and data models, but the changed arrangement prevents direct identification of sensitive information. This allows testing effectiveness to be maintained while reducing disclosure risk.

Inventive Principle:
Principle #35Parameter changes

2Object-affected harmful factors

If artificial data is used for testing, then sensitive information disclosure risk is reduced, but testing effectiveness deteriorates

Engineering Contradiction:
Improvesensitive information disclosure riskVSAvoidtesting effectiveness
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

Rather than using artificially generated data, the patent copies actual production data into infocubes and applies shuffling transformations. This copied and transformed data retains the real-world statistical properties, relationships, and patterns necessary for effective testing of statistical functions and data models, while still protecting sensitive information through the shuffling operation.

Inventive Principle:
Principle #26Copying

3Object-affected harmful factors

If data shuffling is performed on infocube data, then data anonymity is improved, but data structure complexity increases

Engineering Contradiction:
Improvedata anonymityVSAvoiddata structure complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The patent segments the data structure into infocubes with specific organizational patterns (e.g., date-based segmentation, hierarchical segmentation). This structured segmentation allows for systematic shuffling operations that maintain data relationships while achieving anonymity. The segmented structure makes the shuffling process more manageable and the resulting anonymous data more useful for testing than completely random arrangements.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS7779041B2Anonymizing infocube data
Publication Date: 2010.08.17 SAP SE
  • US7779041B2 patent drawing
  • US7779041B2 patent drawing
  • US7779041B2 patent drawing

AI summary

In one aspect, in a computer-implemented method may make data anonymous, so that the data may be used during testing. The method may include receiving, from a user interface, an indication of a type of shuffling to be performed on data. Moreover, the data may be shuffled based on the received indication of the shuffling type. The shuffling may rearrange the data to make the data anonymous. The shuffled data may be provided to an application. Related systems, apparatus, methods, and/or articles are also described.