Wellbore Data Anonymization via Shuffling and Normalization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Downhole exploration and production data anonymization is challenging due to the need to protect confidential information while enabling analysis and sharing for improved wellbore operations, as existing methods fail to effectively remove identifiable data without compromising proprietary information.
Innovation Solution
A computer-implemented method and system for anonymizing data by shuffling, normalizing, and non-dimensionalizing raw data from wellbore operations, allowing for the generation of anonymized data that can be analyzed without revealing sensitive information, and aggregated for performing actions such as drilling, completion, or production decisions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If data anonymization is performed to protect confidential information, then data security is improved, but data utility for analysis may deteriorate
Solution Approach 1:
The patent segments the data anonymization process into distinct operations: shuffling data points to break associations, normalizing data values to standard ranges, and non-dimensionalizing data to remove units. Each operation independently contributes to anonymization while preserving analytical utility, resolving the contradiction between security and usefulness.
Solution Approach 2:
The patent transforms data parameters through multiple conversions: changing the order of data points via shuffling, scaling values through normalization, and removing dimensional units through non-dimensionalization. These parameter changes maintain the essential analytical characteristics while eliminating identifiable information, simultaneously achieving security and utility.
2Quantity of substance
If raw data with depth associations is processed, then data completeness is improved, but data confidentiality deteriorates
Solution Approach 1:
The patent extracts and removes the depth association information from the raw data while preserving the remaining data points for analysis. By separating the identifiable depth information from the measurement data, the system maintains data completeness for analytical purposes while eliminating the confidential identifier.
Solution Approach 2:
Instead of removing data to protect confidentiality, the patent inverts the approach by transforming and repositioning data points through shuffling and normalization. This inversion maintains data completeness while achieving confidentiality through the transformation process itself rather than through data removal.
3Adaptability or versatility
If data is normalized and non-dimensionalized, then data comparability is improved, but data specificity deteriorates
Solution Approach 1:
The patent applies universal transformation operations (normalization and non-dimensionalization) that can be applied across different data types and wellbore operations. These operations create a universal data format that enhances comparability and adaptability across diverse datasets while maintaining the essential specificities needed for analysis through the preservation of data relationships and patterns.
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
Examples of techniques for anonymizing data are disclosed. In one example implementation according to aspects of the present disclosure, a computer-implemented method includes receiving, by a processing device, raw data from a wellbore operation. The raw data can be associated with depths. The method further includes anonymizing, by the processing device, the raw data to convert the raw data to anonymized data. One or more techniques can be implemented to anonymize the data, such as shuffling the raw data, normalizing the raw data, and/or non-dimensionalizing the raw data. The method further includes analyzing, by the processing device, the anonymized data. The method further includes performing an action at the wellbore operation based at least in part on the analysis of the anonymized data.