In-Network Anonymization Engine for Confidential Sentiment Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing de-identification and redaction tools require data to be uploaded to third-party services, posing security risks, are inefficient for large data volumes, and fail to handle domain-specific or unstructured data effectively.
Innovation Solution
A system and method for anonymizing data within an internal network environment, using speech-to-text transcription and domain-specific anonymization techniques to generate anonymized transcripts without relying on third-party APIs, handling large volumes efficiently and accommodating non-uniform data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing de-identification and redaction tools are used, then data anonymization can be performed, but data must be uploaded to third-party services which creates security risks and violates fiduciary duty
Solution Approach 1:
The system performs de-identification and redaction operations locally within the organization's own computing environment rather than requiring external third-party services. The anonymization engine processes data in-place, eliminating the need to upload sensitive customer data outside the secure network boundary while maintaining full anonymization functionality.
Solution Approach 2:
The patent introduces an intermediary anonymization engine that acts as a mediator between the raw data and the analysis process. This intermediary component performs the de-identification and redaction operations within the secure internal environment, preventing direct exposure of sensitive data to external systems while still enabling the required data processing and analysis capabilities.
2Reliability
If third-party redaction services are used, then data can be anonymized, but processing large volumes of data becomes computationally infeasible and expensive
Solution Approach 1:
The system leverages the organization's own computational resources to perform large-scale de-identification and rediction operations locally. By deploying the anonymization engine within the internal computing environment, the system can process terabytes of data without being constrained by third-party service rate limits or per-call API fees, achieving both high productivity and data security.
3Adaptability or versatility
If commercial de-identification tools are used, then general purpose data can be anonymized, but domain-specific data such as insurance policy numbers and medical procedure information cannot be handled
Solution Approach 1:
The anonymization engine features dynamic, configurable redaction patterns that can be adapted to different domains. The system allows users to define custom redaction rules and patterns specific to their domain (such as insurance policy numbers, medical procedures, or other proprietary data formats), enabling the same platform to handle both general and domain-specific data without requiring separate specialized tools.
4Adaptability or versatility
If existing anonymization tools are used, then structured data can be processed, but unstructured and non-uniform data such as various telephone number formats cannot be consistently handled
Solution Approach 1:
The system employs dynamic pattern recognition and adaptive rediction algorithms that automatically detect and handle multiple variations of unstructured data formats. The anonymization engine can identify different telephone number formats, alphanumeric doppelgangers, and other non-uniform data patterns, applying appropriate redaction rules consistently across all variations without requiring manual configuration for each data type.
Data Source
AI summary
A method for anonymizing data includes receiving call data of a call in an interaction recording system located behind a firewall of an internal network sub-environment, and within the internal network sub-environment: (i) storing the call data including interaction metadata, (ii) generating a speech-to-text transcript corresponding to words spoken by one or more callers, and (iii) generating an anonymized transcript by anonymizing personally identifiable information. A computing system includes a processor, and a memory including computer executable instructions that, when executed by the one processor, cause the system to perform the method. A non-transitory computer readable medium contains program instructions that when executed, cause a computer system to perform the method.


