Language Model Pre-Training with Hierarchical Multi-Task Templates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multitask-based finetuning and pre-training technologies for language models lack the ability to learn general knowledge from unsupervised data, leading to limited diversity and robustness in template design.
Innovation Solution
Construct a pre-training language data set comprising unsupervised and supervised data, generate a hierarchical multi-template and multi-task data set, and pre-train the language model using this data set to enhance continuous learning and diversity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multitask-based finetuning technology is used, then the model can perform specific tasks well, but the model cannot learn general knowledge from unsupervised data and cannot learn continuously
Solution Approach 1:
The patent applies preliminary action by performing pre-training on unsupervised data before task-specific finetuning. The language model is first pre-trained on large-scale unsupervised corpora to learn general language knowledge and representations, then subsequently fine-tuned on supervised task-specific data. This preliminary pre-training step enables the model to acquire general knowledge that can be continuously updated and transferred across different tasks.
2Adaptability or versatility
If template design for the model is used, then the model can handle multiple tasks, but the template design lacks diversity which affects the robustness of the model
Solution Approach 1:
The patent applies dimensionality change by transforming the traditional single-template approach into a hierarchical multi-template structure with multiple dimensions. Instead of using one flat template for all tasks, the system creates a hierarchy of templates at different levels (general templates, task-specific templates, and instance-level templates). This multi-dimensional template hierarchy increases design diversity and improves model robustness by providing multiple pathways for handling different task variations.
3Ease of manufacture
If a single pre-training data set is used, then the training process is simple, but the model lacks diversity and robustness in handling different tasks
Solution Approach 1:
The patent applies segmentation by dividing the pre-training process and data set into distinct segments or stages. Instead of using a single undifferentiated data set, the system segments the training data into different types (unsupervised corpora, supervised task data, continuation data) and processes them in sequential stages. This segmentation allows the model to learn different types of knowledge from different data segments, improving task diversity handling while maintaining training organization and simplicity.
Data Source
AI summary
A method for pre-training a language model includes: constructing a pre-training language data set, in which the pre-training language data set comprises unsupervised language data and supervised language data; generating a hierarchical multi-template and multi-task language data set based on the pre-training language data set; and pre-training the language model based on the hierarchical multi-template and multi-task language data set.


