Cross-Referencing Tokens for Structured and Freeform Data Security
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data security technologies, such as encryption and vaultless tokenization, are inadequate for protecting sensitive data in structured and unstructured content due to their vulnerabilities and the need for format preservation, leading to complex enterprise data management and compliance risks.
Innovation Solution
The development of format-preserving and self-describing tokens that can be generated and utilized in structured and unstructured data, allowing for secure data handling and compliance with regulations by replacing sensitive data with tokens that retain format and structure, while enabling secure storage and retrieval of original values.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional tokens are used to protect sensitive data, then data security is improved, but the tokens lack extrinsic meaning and value, restricting their use across various data types
Solution Approach 1:
The patent applies parameter changes by transforming traditional tokens into format-preserving tokens that maintain specific parameters of the original data (length, format, structure) while still providing security. This allows tokens to adapt to different data types and structures, resolving the contradiction between security and versatility by changing the token parameters to preserve both security properties and data characteristics.
Solution Approach 2:
The patent implements universality by creating a tokenization system that can handle multiple data types (structured, unstructured, freeform text) with a single token framework. The format-preserving tokens can be universally applied across different contexts while maintaining their security function, allowing one token system to serve multiple purposes and data formats.
2Reliability
If tokens replace sensitive data values, then data protection is improved, but complex handling is required in complex data environments
Solution Approach 1:
The patent applies self-service through self-describing tokens that contain embedded metadata about their own structure, format, and validation rules. The tokens automatically describe their requirements, enabling systems to handle them without complex external configuration or management overhead, thus reducing system complexity while maintaining protection.
Solution Approach 2:
The patent implements preliminary action by pre-defining token formats, structures, and validation rules before tokens are generated. The format-preserving tokenization process establishes clear templates and patterns in advance, which simplifies subsequent token handling and processing by eliminating the need for complex runtime analysis or configuration.
3Manufacturing precision
If format-preserving tokens are generated for structured data, then data format preservation is improved, but maintaining referential integrity across multiple data types becomes challenging
Solution Approach 1:
The patent uses parameter changes by systematically varying token parameters based on the specific data type and structure being tokenized. Different parameter sets are applied for different data formats (structured, unstructured, freeform text), allowing the system to maintain referential integrity across multiple data types by adapting parameters rather than using a single complex management approach.
Data Source
AI summary
Multiple types of tokens can be generated and utilized in a highly structured document with freeform text. For example, a tokenization system may receive a request for tokenizing a document with a first portion having structured content and a second portion having unstructured or semi-structured content. In response, the tokenization system identifies sensitive information in the first portion of the document, generates format-preserving tokens for the sensitive information in the first portion of the document, identifies sensitive information in the second portion of the document, and generates self-describing tokens for the sensitive information in the second portion of the document. The self-describing tokens reference the sensitive information in the first portion of the document. The tokenization system may then communicate the format-preserving tokens and the self-describing tokens to the first client computing system or to a second client computing system.


