Multilingual Sensitive Data Detection via Neural Text Unification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data search techniques are limited in identifying sensitive information across languages and often require specific phrase matching, failing to prevent accidental transmission of private data.

Innovation Solution

The method involves character, phrase, and concept unification processes to standardize text, allowing for multilingual search and identification of sensitive information, with a processor applying these processes and predefined rules to alert and prevent the transmission of sensitive content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text-based search techniques are used to identify sensitive information, then specific phrase matching can be achieved, but the system fails to identify sensitive information across different languages and requires translations

Engineering Contradiction:
Improveidentification accuracy of sensitive informationVSAvoidmultilingual capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system changes the parameter of text representation by converting different language texts into a unified representation space using neural networks. This allows the same search and identification mechanisms to work across multiple languages without requiring separate language-specific processing, thereby achieving both high identification accuracy and multilingual versatility.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a neural network-based text representation layer that acts as an intermediary between raw multilingual text and the search/query processing layer. This intermediary converts diverse language inputs into a unified representation format, enabling accurate sensitive information identification across languages without requiring direct translation or language-specific processing.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If existing text search techniques are used, then simple phrase matching is possible, but the system cannot identify variations of sensitive information patterns

Engineering Contradiction:
Improvesimplicity of search operationVSAvoiddetection accuracy of sensitive information variations
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent replaces traditional mechanical string-matching algorithms with a neural network-based semantic understanding system. This substitution allows the system to go beyond simple phrase matching and detect variations of sensitive information patterns while maintaining ease of operation through unified search interfaces. The neural networks capture semantic meanings and relationships, enabling accurate detection of sensitive information in various forms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If unification processes are applied to standardize text for multilingual search, then cross-language identification capability is improved, but the processing complexity increases

Engineering Contradiction:
Improvecross-language search capabilityVSAvoidtext processing pipeline complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a universal text representation model that serves multiple functions: it processes texts from different languages, identifies sensitive information, and enables search queries. This multi-functional approach consolidates what would otherwise require separate processing pipelines into a single unified system, achieving cross-language capability while managing processing complexity through shared computational resources and unified architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10984127B2Content leakage protection
Publication Date: 2021.04.20 SOPHOS LTD
  • US10984127B2 patent drawing
  • US10984127B2 patent drawing
  • US10984127B2 patent drawing

AI summary

Methods and systems for identifying content of interest. Accessed textual information is processed by at least one of character unification, phrase unification, and concept unification. A configured processor executes at least one predefined rule to determine whether the unified content includes certain types of information. Unified content that matches may be subject to further action such as alerts, encryption, logging, etc.