Meta-Knowledge Fine-Tuning for Domain-Invariant Language Model Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing language model compression methods are limited by specific data sets, resulting in poor generalization and parameter initialization across different domains for downstream tasks.

Innovation Solution

A method and platform for meta-knowledge fine-tuning based on domain-invariant features, which involves constructing an adversarial domain classifier, combining word and domain embeddings, and using a domain damage objective function to learn features independent of specific domains, enabling better parameter initialization and generalization across different data sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing compression methods fine-tune on specific downstream task data sets, then the model achieves good performance on that specific task, but the model shows poor generalization and parameter initialization ability across different domains

Engineering Contradiction:
Improveperformance on specific taskVSAvoidgeneralization ability across different domains
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by performing meta-knowledge fine-tuning before actual downstream task deployment. The method pre-learns domain-invariant features from multiple source domains using an adversarial domain classifier, so that when the model is later deployed on any target domain, it already possesses transferable knowledge and better parameter initialization, reducing the need for extensive task-specific fine-tuning

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements universality by creating a model that functions across multiple domains simultaneously. The domain-invariant feature learning process enables the model to acquire knowledge that is applicable to various downstream tasks and domains, making a single model architecture universally useful rather than requiring separate fine-tuned models for each specific task

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If domain-specific features are learned from specific data sets, then the model achieves high accuracy on that data set, but the model cannot effectively transfer knowledge to other domains

Engineering Contradiction:
Improveaccuracy on specific data setVSAvoidknowledge transfer capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies the extraction principle by separating domain-specific features from domain-invariant features. The adversarial domain classifier is designed to learn and maximize domain-specific characteristics, while the main model learns domain-invariant features by minimizing the domain classifier's accuracy. This extraction allows the model to discard harmful domain-specific biases while retaining useful general patterns

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements parameter changes by dynamically adjusting the loss function weights and model parameters during the adversarial training process. The method modifies the optimization objectives to balance between maintaining accuracy on source domains and achieving invariance across domains, thereby transforming the model's parameter space to achieve both precision and transferability

Inventive Principle:
Principle #35Parameter changes

3Productivity

If a universal compression architecture is implemented, then the model can be deployed in resource-constrained environments, but the model may lose domain-specific performance optimizations

Engineering Contradiction:
Improvedeployment efficiency in resource-constrained environmentsVSAvoiddomain-specific performance
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies local quality by allowing different parts of the model to have different properties. The universal compression architecture maintains a standardized core structure for efficiency, while domain-specific adaptations can be applied locally through the learned domain-invariant features and selective fine-tuning of specific model components, ensuring both deployment efficiency and domain performance

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11669741B2Method and platform for meta-knowledge fine-tuning based on domain-invariant features
Publication Date: 2023.06.06 ZHEJIANG LAB
  • US11669741B2 patent drawing

AI summary

Disclosed is a method for meta-knowledge fine-tuning and platform based on domain-invariant features. According to the method, highly transferable common knowledge, i.e., domain-invariant features, in different data sets of the same kind of tasks is learnt, the common domain features in different domains corresponding to different data sets of the same kind of tasks learnt in the network set are fine-tuned to be quickly adapted to any different domains. According to the present application, the parameter initialization ability and generalization ability of the universal language model of the same kind of tasks are improved, and finally a common compression framework of the universal language model of the same kind of downstream tasks is obtained through fine tuning. In the meta-knowledge fine-tuning network, a loss function of the domain-invariant features is designed in the present application, and domain-independent universal knowledge is learn.