Language Model Pre-Training with Hierarchical Multi-Task Templates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multitask-based finetuning and pre-training technologies for language models lack the ability to learn general knowledge from unsupervised data, leading to limited diversity and robustness in template design.

Innovation Solution

Construct a pre-training language data set comprising unsupervised and supervised data, generate a hierarchical multi-template and multi-task data set, and pre-train the language model using this data set to enhance continuous learning and diversity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multitask-based finetuning technology is used, then the model can perform specific tasks well, but the model cannot learn general knowledge from unsupervised data and cannot learn continuously

Engineering Contradiction:
Improvetask performanceVSAvoidcontinuous learning capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by performing pre-training on unsupervised data before task-specific finetuning. The language model is first pre-trained on large-scale unsupervised corpora to learn general language knowledge and representations, then subsequently fine-tuned on supervised task-specific data. This preliminary pre-training step enables the model to acquire general knowledge that can be continuously updated and transferred across different tasks.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If template design for the model is used, then the model can handle multiple tasks, but the template design lacks diversity which affects the robustness of the model

Engineering Contradiction:
Improvemulti-task capabilityVSAvoidmodel robustness
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies dimensionality change by transforming the traditional single-template approach into a hierarchical multi-template structure with multiple dimensions. Instead of using one flat template for all tasks, the system creates a hierarchy of templates at different levels (general templates, task-specific templates, and instance-level templates). This multi-dimensional template hierarchy increases design diversity and improves model robustness by providing multiple pathways for handling different task variations.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Ease of manufacture

If a single pre-training data set is used, then the training process is simple, but the model lacks diversity and robustness in handling different tasks

Engineering Contradiction:
Improvetraining simplicityVSAvoidtask diversity handling
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent applies segmentation by dividing the pre-training process and data set into distinct segments or stages. Instead of using a single undifferentiated data set, the system segments the training data into different types (unsupervised corpora, supervised task data, continuation data) and processes them in sequential stages. This segmentation allows the model to learn different types of knowledge from different data segments, improving task diversity handling while maintaining training organization and simplicity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12373735B2Method for pre-training language model
Publication Date: 2025.07.29 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12373735B2 patent drawing
  • US12373735B2 patent drawing
  • US12373735B2 patent drawing

AI summary

A method for pre-training a language model includes: constructing a pre-training language data set, in which the pre-training language data set comprises unsupervised language data and supervised language data; generating a hierarchical multi-template and multi-task language data set based on the pre-training language data set; and pre-training the language model based on the hierarchical multi-template and multi-task language data set.