Multilingual Sensitive Data Detection via Neural Text Unification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data search techniques are limited in identifying sensitive information across languages and often require specific phrase matching, failing to prevent accidental transmission of private data.
Innovation Solution
The method involves character, phrase, and concept unification processes to standardize text, allowing for multilingual search and identification of sensitive information, with a processor applying these processes and predefined rules to alert and prevent the transmission of sensitive content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If text-based search techniques are used to identify sensitive information, then specific phrase matching can be achieved, but the system fails to identify sensitive information across different languages and requires translations
Solution Approach 1:
The system changes the parameter of text representation by converting different language texts into a unified representation space using neural networks. This allows the same search and identification mechanisms to work across multiple languages without requiring separate language-specific processing, thereby achieving both high identification accuracy and multilingual versatility.
Solution Approach 2:
The patent introduces a neural network-based text representation layer that acts as an intermediary between raw multilingual text and the search/query processing layer. This intermediary converts diverse language inputs into a unified representation format, enabling accurate sensitive information identification across languages without requiring direct translation or language-specific processing.
2Ease of operation
If existing text search techniques are used, then simple phrase matching is possible, but the system cannot identify variations of sensitive information patterns
Solution Approach 1:
The patent replaces traditional mechanical string-matching algorithms with a neural network-based semantic understanding system. This substitution allows the system to go beyond simple phrase matching and detect variations of sensitive information patterns while maintaining ease of operation through unified search interfaces. The neural networks capture semantic meanings and relationships, enabling accurate detection of sensitive information in various forms.
3Adaptability or versatility
If unification processes are applied to standardize text for multilingual search, then cross-language identification capability is improved, but the processing complexity increases
Solution Approach 1:
The patent implements a universal text representation model that serves multiple functions: it processes texts from different languages, identifies sensitive information, and enables search queries. This multi-functional approach consolidates what would otherwise require separate processing pipelines into a single unified system, achieving cross-language capability while managing processing complexity through shared computational resources and unified architecture.
Data Source
AI summary
Methods and systems for identifying content of interest. Accessed textual information is processed by at least one of character unification, phrase unification, and concept unification. A configured processor executes at least one predefined rule to determine whether the unified content includes certain types of information. Unified content that matches may be subject to further action such as alerts, encryption, logging, etc.


