Infocube Data Shuffling for Secure Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In the context of SAP's Business Warehouse system, using actual production data for testing poses a risk of sensitive information disclosure, and relying on artificial data limits the effectiveness of testing, necessitating a mechanism to create anonymous data for secure testing purposes.
Innovation Solution
A method and apparatus for shuffling data within infocubes, specifically rearranging columns of fact, dimension, and characteristic tables using random number assignment or dictionaries, to make the data anonymous and protect sensitive information, enabling its use in testing applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If actual production data is used for testing, then testing effectiveness is improved, but risk of sensitive information disclosure increases
Solution Approach 1:
The patent creates a copy of the actual production data through infocubes, then applies shuffling operations to this copy. The shuffled infocube data maintains the statistical properties and relationships of the original data for effective testing, while the values are permuted to prevent identification of sensitive information. This resolves the contradiction by using a transformed copy rather than the original data.
Solution Approach 2:
The patent transforms the data by changing the parameter arrangement through shuffling operations on infocube values. By permuting the values while maintaining their distribution characteristics, the data remains useful for testing statistical functions and data models, but the changed arrangement prevents direct identification of sensitive information. This allows testing effectiveness to be maintained while reducing disclosure risk.
2Object-affected harmful factors
If artificial data is used for testing, then sensitive information disclosure risk is reduced, but testing effectiveness deteriorates
Solution Approach 1:
Rather than using artificially generated data, the patent copies actual production data into infocubes and applies shuffling transformations. This copied and transformed data retains the real-world statistical properties, relationships, and patterns necessary for effective testing of statistical functions and data models, while still protecting sensitive information through the shuffling operation.
3Object-affected harmful factors
If data shuffling is performed on infocube data, then data anonymity is improved, but data structure complexity increases
Solution Approach 1:
The patent segments the data structure into infocubes with specific organizational patterns (e.g., date-based segmentation, hierarchical segmentation). This structured segmentation allows for systematic shuffling operations that maintain data relationships while achieving anonymity. The segmented structure makes the shuffling process more manageable and the resulting anonymous data more useful for testing than completely random arrangements.
Data Source
AI summary
In one aspect, in a computer-implemented method may make data anonymous, so that the data may be used during testing. The method may include receiving, from a user interface, an indication of a type of shuffling to be performed on data. Moreover, the data may be shuffled based on the received indication of the shuffling type. The shuffling may rearrange the data to make the data anonymous. The shuffled data may be provided to an application. Related systems, apparatus, methods, and/or articles are also described.


