Federated Tensor Text Analysis for Confidential Anomaly Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing document management systems fail to provide a convenient mechanism for objective and subjective comparison of documents, particularly in understanding how a document's text compares to similar text within other documents, which is crucial for analyzing legal documents like contracts, leading to costly and time-consuming efforts to assess termination clauses and their compliance with industry standards.
Innovation Solution
A federated system that converts text portions into tensors and compares them across multiple computing environments, calculating similarity scores without accessing the underlying documents, using pre-defined hashing algorithms, trained vocabulary-based methods, or machine learning models, while maintaining document confidentiality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional document management systems are used to compare documents, then document organization and version control are provided, but objective and subjective comparison of documents across different sources is not available
Solution Approach 1:
The patent introduces tensor representations as an intermediary mechanism that transforms text portions into numerical formats suitable for mathematical comparison. This intermediary enables objective comparison across different document sources while maintaining the original document structures intact, resolving the contradiction between measurement precision and ease of operation.
Solution Approach 2:
The patent replaces manual or simple textual comparison mechanisms with tensor-based mathematical operations. By substituting the mechanical/textual comparison process with tensor algebra operations, the system achieves precise objective comparison while providing automated convenience, addressing both measurement precision and ease of operation requirements.
2Measurement precision
If manual review of documents is performed to assess text accuracy and conformity, then detailed analysis is achieved, but significant time and cost are required
Solution Approach 1:
The patent substitutes manual text review with automated tensor-based comparison systems. By representing text portions as tensors and applying mathematical operations, the system achieves detailed text analysis accuracy while eliminating the time-consuming nature of manual review, directly addressing the contradiction between measurement precision and time loss.
Solution Approach 2:
The patent transforms text from its original linguistic form into tensor parameters that enable quantitative comparison. This parameter change allows computational processing of text data, achieving accurate analysis results while dramatically reducing the time required compared to manual review methods.
3Loss of information
If documents are accessed and compared across multiple sources, then comprehensive comparison is achieved, but document confidentiality and security are compromised
Solution Approach 1:
The patent extracts only the essential text portions needed for comparison and converts them into tensors, leaving the original confidential documents intact and inaccessible to external systems. This extraction approach enables comprehensive comparison information to be obtained while the sensitive source documents remain secure, resolving the contradiction between information completeness and security protection.
Solution Approach 2:
The patent uses tensor representations as an intermediary that decouples the comparison process from the original confidential documents. By working with tensor transformations rather than direct document access, the system achieves complete comparison information while maintaining document confidentiality, as the tensors can be processed without exposing the underlying sensitive data.
Data Source
AI summary
Aspects of the present disclosure involve systems and methods for evaluating a piece of text or document against many corpuses of text or documents located on sources which may be the same and/or different from the text of interest in a tensorized manner and aggregating the coherence/anomaly score against some or all of the entire corpus. This joining of multiple data sources for evaluating the given piece or text may be a “federated” system as disparate data sources, each of which may contain confidential or otherwise private information, may be considered as a single repository of texts or documents. The systems and methods provide for a coherency and/or anomaly check of a piece of text of a document against similar pieces of text to determine a similarity of the piece of text to a large corpus of documents stored in disparate locations.


