Personal Data Classification System Using Graph Linkage

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data protection methods are inadequate for accurately extracting, classifying, and linking personal data entities to comply with stringent regulations like GDPR, leading to potential fines and data privacy breaches.

Innovation Solution

A system and method for personal data classification that includes entity extraction, linkage, and purpose prediction using graph-based methodologies, deep learning, and machine learning models for identifying relationships and predicting the purpose of processing personal data, enabling accurate extraction, classification, and linking of personal data entities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional security and privacy techniques are used, then implementation is simple, but they are overstretched and cannot accurately extract and link personal data entities to comply with stringent regulations

Engineering Contradiction:
Improveaccuracy of personal data extraction and linkingVSAvoidcomplexity of data processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the complex task of personal data protection into distinct functional modules: entity extraction module for identifying personal data entities, linkage module for linking entities to individuals, and purpose prediction module for determining processing purposes. This segmentation enables each module to specialize in specific functions, improving overall accuracy while managing complexity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary knowledge graph that serves as a mediator between raw personal data and compliance requirements. The knowledge graph structures extracted entities and their relationships, enabling accurate linking of personal data to individuals while providing a standardized interface for purpose prediction and regulatory compliance checking.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If advanced deep learning and machine learning models are implemented for accurate personal data extraction and linking, then compliance accuracy improves, but computational resources and processing time increase

Engineering Contradiction:
Improvecompliance with data protection regulationsVSAvoidprocessing time for data extraction and linking
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-processing documents to extract entities and build the knowledge graph before actual compliance checking. The entity extraction module identifies and structures personal data entities in advance, creating a ready-to-use knowledge representation that accelerates subsequent linking and purpose prediction operations, reducing overall processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The purpose prediction module employs self-service mechanisms by automatically inferring processing purposes from contextual features and document content without requiring manual annotation. The system uses machine learning models to autonomously determine purposes such as 'contract fulfillment' or 'marketing' based on extracted entities and document patterns, significantly reducing processing time compared to manual classification.

Inventive Principle:
Principle #25Self-service

3Loss of information

If comprehensive entity extraction and linking is performed across all data repositories, then complete personal data inventory is achieved, but system complexity and resource requirements increase

Engineering Contradiction:
Improvecompleteness of personal data inventoryVSAvoidcomplexity of data repository integration
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The knowledge graph serves as a universal data structure that can represent various types of personal data entities (individuals, organizations, locations) and their relationships across different data repositories. This multi-functional representation enables the system to handle diverse data sources uniformly, achieving complete personal data inventory without proportionally increasing system complexity through standardized integration interfaces.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12039074B2Methods, personal data analysis system for sensitive personal information detection, linking and purposes of personal data usage prediction
Publication Date: 2024.07.16 DATHENA SCI PTE LTD
  • US12039074B2 patent drawing
  • US12039074B2 patent drawing
  • US12039074B2 patent drawing

AI summary

Systems and methods for personal data classification, linkage and purpose of processing prediction are provided. The system for personal data classification includes an entity extraction module for extracting personal data from one or more data repositories in a computer network or cloud infrastructure, a linkage module coupled to the entity extraction module, a linkage module coupled to the entity extraction module and a processing prediction module. The entity extraction module performs entity recognition from the structured, semi-structured and unstructured records in the one or more data repositories. The linkage module uses graph-based methodology to link the personal data to one or more individuals. And the purpose prediction module includes a feature extraction module a purpose of processing prediction module, wherein the feature extraction module extracts both context features and record's features from records in the one or more data repositories, and the purpose of processing prediction module predicts a unique or multiple purpose of processing of the personal data.