Document Marking Using ML for Tone-Preserving Source Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional document marking techniques are susceptible to subversion, such as removal of spacing variations and font changes, and suffer from factual accuracy errors due to lack of consideration for sentence semantics and tone, particularly in fields like corporate and legal documents.
Innovation Solution
The use of generative machine learning models, fine-tuned on a specific document corpus, to modify document terminology while preserving tone and semantics, generating unique copies that are resistant to subversion and maintain factual accuracy, with additional models for tone identification and integrity checks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional document marking techniques are used, then document source detection is implemented, but the documents are susceptible to subversion such as removal of spacing variations and font changes
Solution Approach 1:
The patent changes the parameter of document modification from superficial formatting (spacing, font) to semantic content (synonym replacement). This resolves the contradiction because semantic changes are harder to detect and reverse compared to formatting changes, while still enabling unique identification of document copies through controlled vocabulary substitution.
Solution Approach 2:
The patent introduces an intermediary layer of alternative terminology between the original document and the marked copy. This intermediary vocabulary layer preserves the original meaning while creating detectable differences that are resistant to subversion, as the semantic relationship maintains document integrity while the specific word choices provide unique identifiers.
2Productivity
If machine learning models modify document terminology, then unique copies are generated, but there is a risk of factual accuracy errors due to lack of consideration for sentence semantics and tone
Solution Approach 1:
The patent implements feedback mechanisms where the machine learning model's synonym suggestions are evaluated against multiple criteria including semantic equivalence, tone consistency, and factual accuracy. This feedback loop ensures that only appropriate substitutions are made, resolving the contradiction by maintaining high productivity through automation while ensuring manufacturing precision through multi-layered validation.
Solution Approach 2:
The patent performs preliminary actions by pre-training the machine learning model on domain-specific corpora and establishing guidelines for acceptable synonym substitutions before actual document processing. This preliminary preparation ensures that the model understands sentence semantics and tone requirements in advance, preventing factual accuracy errors while maintaining efficient automated processing.
3Reliability
If multiple machine learning models are used for tone identification and term selection, then tone preservation is improved, but device complexity increases
Solution Approach 1:
The patent segments the complex task of tone-preserving synonym selection into multiple specialized machine learning models, each handling a specific aspect (tone identification, semantic equivalence, factual accuracy). This segmentation resolves the contradiction by improving reliability through specialized functionality while managing device complexity through modular architecture that allows independent training and deployment of each model component.
Data Source
AI summary
Unique copies of an original document can be generated and provided to individual recipients. The unique copies can be used to identify the source of a document leak. The unique copies are generated by replacing terms within the original document with alternative terms. The alternative terms are determined using a first machine learning model that receives a term from the document and outputs the alternative terms. The output alternative terms are provided to a second machine learning model that indicates a tone for each alternative term. The tone of the alternative terms is compared to the tone of the term from the original document, and one or more of the alternative terms are selected based on the tone of the alternative terms relative to the tone of the document term. The alternative terms used to generate the unique copies have a same or similar tone as the document term.


