Joint Pre-training of Multi-modal Question-Answering Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Pre-trained models for question-answering tasks are typically specific to individual forms and require separate training, leading to high resource consumption and time costs, with poor performance in forms like video question-answering due to insufficient training samples.

Innovation Solution

A method for jointly pre-training a model across multiple question-answering forms, including text, knowledge-based, table, image, and video tasks, within a unified framework, allowing knowledge transfer and improving performance across different forms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate pre-trained models are trained for different question-answering forms, then each model can be optimized for its specific form, but resource consumption and time costs increase significantly

Engineering Contradiction:
Improvequestion-answering performanceVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent combines multiple question-answering tasks (text-based, knowledge-based, table-based, image-based, and video-based QA) into a single unified pre-training framework. The model simultaneously learns from diverse QA forms during pre-training, enabling it to handle different QA types without requiring separate models, thus reducing resource consumption and training time while maintaining performance.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a universal pre-trained model that can perform multiple question-answering functions across different modalities. The model is designed with multi-functionality to handle text, knowledge, tables, images, and videos within a single framework, eliminating the need for separate specialized models and achieving both optimization and efficiency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If separate pre-trained models are trained for different question-answering forms, then each model can achieve specialized performance, but the complexity of the system increases

Engineering Contradiction:
Improvequestion-answering performanceVSAvoidmodel management complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple specialized models into a single unified model that handles text-based, knowledge-based, table-based, image-based, and video-based question answering. This consolidation reduces system complexity by eliminating the need to manage multiple separate models while maintaining specialized performance through multi-task learning during pre-training.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If separate pre-trained models are trained for different question-answering forms, then each model can be optimized independently, but time costs and training resources are consumed excessively

Engineering Contradiction:
Improvemodel optimizationVSAvoidpre-training time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary joint pre-training on diverse question-answering tasks before fine-tuning for specific applications. By pre-training the model simultaneously on multiple QA forms (text, knowledge, tables, images, videos) in advance, the model acquires generalized capabilities that reduce the time needed for subsequent specialized training and deployment.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4123516A1Method and apparatus for acquiring pre-trained model, electronic device and storage medium
Publication Date: 2023.01.25 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • EP4123516A1 patent drawingFigure 1
  • EP4123516A1 patent drawingFigure 2~3
  • EP4123516A1 patent drawingFigure 4

AI summary

The present disclosure provides a method and apparatus for acquiring a pre-trained model, an electronic device and a storage medium, and relates to the fields such as deep learning, natural language processing, knowledge graph and intelligent voice. The method may include: acquiring a pre-training task set composed of M pre-training tasks, M being a positive integer greater than 1, the pre-training tasks including: N question-answering tasks corresponding to different question-answering forms, N being a positive integer greater than 1 and less than or equal to M; and jointly pre-training the pre-trained model according to the M pre-training tasks. By use of the solutions of the present disclosure, resource consumption may be reduced, and time costs may be saved.