Set Neural Network Information Loss Reduction via Token Divergence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Set neural networks experience information loss during encoding due to divergence between virtual tokens and input tokens, leading to suboptimal performance and increased training time.

Innovation Solution

Incorporating an information loss term that minimizes the divergence between the distributions of virtual tokens and input tokens using metrics like Kullback-Leibler divergence, allowing for improved training and generalization of set neural networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If set neural networks use virtual tokens for encoding, then the model can process set data, but information loss occurs due to divergence between virtual tokens and input tokens

Engineering Contradiction:
Improveset data processing capabilityVSAvoidinformation loss during encoding
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent introduces an information loss term that computes the divergence between virtual tokens and input tokens, creating a feedback mechanism during training. This feedback guides the optimization process to minimize information loss while maintaining the virtual token encoding approach for set data processing.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent modifies the training objective by adding an information loss term to the loss function. This parameter change adjusts the optimization target to simultaneously achieve set data processing capability and minimize information loss by controlling the divergence between token distributions.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If standard training is used without information loss term, then training is simpler, but convergence is slower and performance is suboptimal

Engineering Contradiction:
Improvetraining simplicityVSAvoidtraining convergence speed
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The information loss term provides continuous feedback during training about the divergence between virtual and input tokens. This feedback accelerates convergence by guiding the model to learn more effective token representations, improving training productivity without significantly complicating the training procedure.

Inventive Principle:
Principle #23Feedback

3Reliability

If more training iterations are performed to reduce information loss, then model performance improves, but computational resources and time increase

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The information loss term performs preliminary action by pre-aligning the virtual token distribution with the input token distribution during training. This preliminary alignment prevents information loss accumulation, allowing the model to achieve high performance with fewer training iterations and reduced computational resources.

Inventive Principle:
Principle #10Preliminary action

4Device complexity

If virtual tokens are used for encoding, then the model structure is simplified, but information divergence from input tokens occurs

Engineering Contradiction:
Improvemodel structure complexityVSAvoidtoken distribution divergence
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent implements a feedback mechanism through the information loss term that monitors and minimizes the divergence between virtual token and input token distributions. This feedback ensures that the simplified model structure using virtual tokens does not sacrifice information fidelity.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

By modifying the loss function to include an information loss term, the patent changes the optimization parameters to simultaneously achieve structural simplification through virtual tokens and maintain information fidelity by controlling token distribution divergence.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20230077692A1Mechanism for reducing information lost in set neural networks
Publication Date: 2023.03.16 NEC LAB EURO GMBH
  • US20230077692A1 patent drawing
  • US20230077692A1 patent drawing
  • US20230077692A1 patent drawing

AI summary

A method for minimizing information loss in set neural networks includes determining an information loss term for a set neural network that internally uses virtual tokens, such that the information loss term minimizes a divergence between two distributions. The set neural network is trained with training data from a data source that is expressed as sets using the information loss term.