Neural Network Training via Divergent Probe Inputs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Training an accurate neural network is complex and time-consuming, and existing methods require access to the original training dataset to modify or replicate a pre-trained neural network, making it difficult to incorporate new information or improve the model without overriding existing knowledge.

Innovation Solution

A method to train a new neural network to mimic a pre-trained target network by generating a divergent probe training dataset that maximizes differences between the student and mentor networks' outputs, allowing for efficient training and modification without access to the original training data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional training methods are used to train an accurate neural network, then the model achieves high accuracy, but the training process becomes extremely time-consuming and computationally intensive

Engineering Contradiction:
Improvemodel accuracyVSAvoidtraining time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent extracts only the essential training information needed from the pre-trained model by using probe inputs to generate outputs that capture the model's behavior patterns. Instead of retraining on the entire original dataset, the system extracts representative input-output pairs that are sufficient for training student models, significantly reducing training time while maintaining accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameters of the training approach by using divergent probe inputs that maximize output differences between student and mentor models. This parameter change in the training methodology allows for faster convergence while achieving the same level of model accuracy, resolving the contradiction between training time and model precision.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If the original training dataset is kept secret for data privacy and security reasons, then data security is maintained, but other devices cannot replicate or modify the pre-trained neural network

Engineering Contradiction:
Improvedata securityVSAvoidmodel replication capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary mechanism - probe inputs and outputs - that allows student models to learn from the pre-trained model without direct access to the secret training data. This intermediary approach enables model replication and modification while maintaining data security, as the probe-based training captures essential patterns without exposing proprietary data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent enables copying of the pre-trained model's behavior and knowledge through probe-based training. Student models can replicate the mentor model's functionality by learning from probe input-output pairs, allowing model distribution and adaptation without sharing the original training dataset, thus maintaining security while enabling versatility.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If the neural network is retrained with new data to incorporate new information, then the model can be updated, but the entire training process must be re-run from scratch which is time-consuming

Engineering Contradiction:
Improvemodel update capabilityVSAvoidretraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by using probe inputs to generate training data that captures the essential patterns of the pre-trained model. This preliminary probe-based training creates a foundation that can be efficiently updated with new data, avoiding the need to re-run the entire training process from scratch while still incorporating new information effectively.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces dynamics to the training process by enabling flexible updates to student models through probe-based learning. The system can dynamically incorporate new data and modify models without static retraining constraints, allowing for efficient model adaptation and updates while maintaining the benefits of the original pre-trained knowledge.

Inventive Principle:
Principle #15Dynamics

4Productivity

If a divergent probe training dataset is used to maximize differences between student and mentor networks, then training is accelerated and convergence is faster, but the training process becomes more complex

Engineering Contradiction:
Improvetraining speedVSAvoidtraining process complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent uses feedback mechanisms by measuring output differences between student and mentor models and using this information to generate divergent probe inputs. This feedback-driven approach systematically identifies and targets the most informative training examples, accelerating convergence while managing complexity through automated feedback loops rather than manual intervention.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent implements self-service by enabling the training system to automatically generate its own probe training data based on measured differences between models. The system serves itself by identifying knowledge gaps and autonomously creating targeted training examples, which accelerates training while reducing the complexity burden on external systems or operators.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240220807A1Training a student neural network to mimic a mentor neural network with inputs that maximize student-to-mentor disagreement
Publication Date: 2024.07.04 NANO DIMENSIONS TECH LTD
  • US20240220807A1 patent drawing
  • US20240220807A1 patent drawing
  • US20240220807A1 patent drawing

AI summary

A device, system, and method is provided for training a new neural network to mimic a target neural network without access to the target neural network or its original training dataset. The target neural network and the new neural network may be probed with input data to generate corresponding target and new output data. Input data may be detected that generate a maximum or above threshold difference between the corresponding target and new output data. A divergent probe training dataset may be generated comprising the input data that generate the maximum or above threshold difference and the corresponding target output data. The new neural network may be trained using the divergent probe training dataset to generate the target output data. The new neural network may be iteratively trained using an updated divergent probe training dataset dynamically adjusted as the new neural network changes during training.