Document Entity Detection via Clause Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing database systems fail to efficiently preserve and reuse semantic information from physical documents, and struggle to identify and manage personal or potentially-personal data entities as required by regulations like GDPR.
Innovation Solution
A system that scans and digitizes physical documents, extracts and classifies clauses, identifies data privacy and protection entities, and generates weighted associations between documents and entities, using machine learning models and visualization tools to facilitate efficient data management and compliance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If physical documents are scanned and stored into database systems, then electronic access to document content is achieved, but semantic information and substantive information embodied within the documents are not preserved
Solution Approach 1:
The patent segments documents into meaningful portions such as clauses, sections, and paragraphs, and further identifies entities within these portions. This segmentation allows the system to preserve semantic information by organizing it in a structured manner that maintains the meaning and context of different document elements, rather than treating the document as a single unstructured block of text.
Solution Approach 2:
The patent introduces an intermediary layer between the scanned document image and the database storage. This intermediary consists of optical character recognition (OCR) technology and natural language processing (NLP) components that extract and structure semantic information from the scanned document, transforming it into a format that preserves meaning while enabling electronic access.
2Stability of the object's composition
If documents are stored as complete units, then document integrity is maintained, but independent re-use of document portions is not enabled
Solution Approach 1:
The patent divides documents into hierarchical segments including clauses, sections, paragraphs, and entities. Each segment is stored with its contextual relationships to other segments, allowing the system to maintain document integrity through structured relationships while enabling flexible re-use of individual portions. Users can retrieve and reuse specific clauses or entities independently while the system preserves the overall document structure.
Solution Approach 2:
The patent implements a nested structure where entities are nested within clauses, clauses are nested within sections, and sections are nested within the complete document. This hierarchical nesting allows the system to maintain the integrity of the complete document while enabling access and re-use of nested portions at any level, as each level preserves its contextual relationships with parent and child elements.
3Productivity
If traditional scanning methods are used, then document digitization is achieved, but identification of data privacy entities and determination of privacy-weighted associations are not possible
Solution Approach 1:
The patent introduces NLP-based entity recognition components as intermediaries between the OCR-extracted text and the final database storage. These intermediaries automatically identify data privacy entities such as personal names, organization names, locations, and other sensitive information, and determine privacy-weighted associations between entities and document portions, enabling GDPR compliance without manual review.
Solution Approach 2:
The patent implements automated entity recognition and classification systems that perform data privacy analysis without human intervention. The system automatically scans document text, identifies sensitive entities, classifies them by type and privacy sensitivity, and determines associations with specific document portions, enabling the system to serve its own compliance needs without requiring external expert analysis.
Data Source
AI summary
Systems and methods include extraction of a plurality of clauses from each of a plurality of electronic documents, determination, for each of the plurality of clauses and using a machine-learned algorithm, an associated clause type, identification of one or more data privacy protection entities present within each of one or more of the plurality of clauses, determination, for each of the one or more of the plurality of clauses, of a weighted frequency for each of the one or more data privacy protection entities present within the clause based on a type of the data privacy protection entity, determination of a weighted frequency associated with each of the plurality of electronic documents based on the determined weighted frequency for each of the one or more data privacy protection entities present within clauses of the plurality of electronic documents, and storage of an identifier of each of the plurality of electronic documents in association with a respective determined weighted frequency.


