Distributed Ledger for Secure Model Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional approaches struggle to efficiently track and manage frequent changes in contracts, amendments, and addendums, particularly due to their focus on unstructured text and inability to extract relevant information correctly, leading to risks associated with overlooking material contractual terminologies.

Innovation Solution

A distributed machine learning model is generated and trained using data from multiple third parties via a distributed ledger system, such as a blockchain, allowing each party to control and track their data submissions and withdrawals, while maintaining a record of data contributions and model usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If data from multiple third parties is collected for model training, then the comprehensiveness and accuracy of the machine learning model is improved, but data security and privacy protection become more difficult to ensure

Engineering Contradiction:
Improvemodel training accuracyVSAvoiddata security risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the data storage and model training process into distributed components across multiple blockchain nodes. Each third party's data is stored in encrypted form on their own device or a distributed file system, while only hashed references are stored on the blockchain. The model training is performed in a decentralized manner where multiple parties contribute computational resources and data samples without exposing their complete datasets, thus improving model comprehensiveness while maintaining data security through segmentation of data ownership and control.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces blockchain technology as an intermediary layer between data contributors and the model training process. The blockchain ledger serves as a trusted mediator that records data submission proofs, manages access permissions, and coordinates the distributed training process without any single party having access to all raw data. This intermediary mechanism enables multiple parties to collaborate on model training while maintaining data privacy and security through cryptographic verification and decentralized coordination.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If conventional approaches focus on unstructured text for contract analysis, then the simplicity of the system is maintained, but the ability to extract relevant information correctly deteriorates

Engineering Contradiction:
Improvesystem simplicityVSAvoidinformation extraction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent replaces conventional mechanical text processing approaches with machine learning models that use natural language processing and semantic analysis. Instead of relying on simple keyword matching or rule-based systems, the patent employs trained machine learning models that can understand contract semantics, identify relevant clauses, and extract meaningful information with high accuracy. These models are trained on diverse contract data from multiple third parties, enabling them to handle the complexity of unstructured legal text while maintaining system scalability through automated processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of operation

If data is stored centrally for model training, then the ease of data access is improved, but the ability of parties to control and remove their data deteriorates

Engineering Contradiction:
Improvedata access convenienceVSAvoiddata control flexibility
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic data management system where data access permissions and participation status can change over time. Third parties can dynamically join or leave the model training process, and their data can be added or removed from the training set through smart contract executions. The system maintains centralized coordination for model training while allowing individual parties to dynamically control their data participation, combining the benefits of easy data access with flexible data control through programmable, time-varying access policies.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250124311A1Managing information for model training using distributed blockchain ledger
Publication Date: 2025.04.17 DOCUSIGN INT EMEA LTD
  • US20250124311A1 patent drawing
  • US20250124311A1 patent drawing
  • US20250124311A1 patent drawing

AI summary

Embodiments are directed to generating and training a distributed machine learning model using data received from a plurality of third parties using a distributed ledger system, such as a blockchain. As each third party submits data suitable for model training, the data submissions are recorded onto the distributed ledger. By traversing the ledger, the learning platform identifies what data has been submitted and by which parties, and trains a model using the submitted data. Each party is also able to remove their data from the learning platform, which is also reflected in the distributed ledger. The distributed ledger thus maintains a record of which parties submitted data, and which parties removed their data from the learning platform, allowing for different third parties to contribute data for model training, while retaining control over their submitted data by being able to remove their data from the learning platform.