Foundation Model Attribution via Trained Attribution Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The proliferation of fine-tuned machine learning models has made it difficult to track their provenance and intellectual property usage, leading to challenges in ensuring model and data transparency and compliance with licensing agreements and laws.

Innovation Solution

A method and system for fine-tuned model to source foundation model attribution, which involves generating training prompt responses, training an attribution model using these responses, and attributing fine-tuned models to their source foundation models based on the trained attribution model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If fine-tuned models are proliferated using existing foundation models, then model versatility and productivity are improved, but model provenance tracking and intellectual property transparency deteriorate

Engineering Contradiction:
Improvemodel proliferation rateVSAvoidmodel provenance information
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent applies preliminary action by embedding attribution information (watermarks, metadata, or fingerprints) into foundation models before they are used to create fine-tuned models. This proactive approach ensures that provenance information is preserved throughout the model lifecycle, enabling automatic tracking and attribution without requiring subsequent intervention or manual record-keeping.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If models are treated as black boxes without access to internal information, then model security and simplicity are maintained, but attribution accuracy and transparency deteriorate

Engineering Contradiction:
Improvemodel accessibilityVSAvoidattribution accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary attribution model that acts as a mediator between the black-box fine-tuned models and the attribution analysis system. This intermediary model processes model responses and extracts attribution signals without requiring direct access to the internal structures of the fine-tuned models, thus maintaining their black-box nature while enabling accurate attribution through indirect observation of model behavior patterns.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If attribution models are trained using generated responses, then attribution capability is improved, but training data requirements and computational resources increase

Engineering Contradiction:
Improveattribution capabilityVSAvoidtraining data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies self-service by having the foundation model generate its own training data through self-attention mechanisms and internal feature analysis. The model creates attribution training examples by processing its own responses and identifying characteristic patterns, eliminating the need for external manual annotation or large-scale data collection. This self-generated training data is then used to train the attribution model, reducing external resource requirements while maintaining high attribution reliability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250028992A1Fine-tuned model to source foundation model attribution
Publication Date: 2025.01.23 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20250028992A1 patent drawing
  • US20250028992A1 patent drawing
  • US20250028992A1 patent drawing

AI summary

An embodiment causes generating, by a trained model, a training prompt response to a training prompt in a set of training prompts. An embodiment trains, using the training prompt and the training prompt response, an attribution model, the training resulting in a trained attribution model. An embodiment attributes, using the trained attribution model and a first prompt response generated by a fine-tuned model in response to a prompt, the fine-tuned model to a foundation model.