Joint Pre-training of Multi-modal Question-Answering Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Pre-trained models for question-answering tasks are typically specific to individual forms and require separate training, leading to high resource consumption and time costs, with poor performance in forms like video question-answering due to insufficient training samples.
Innovation Solution
A method for jointly pre-training a model across multiple question-answering forms, including text, knowledge-based, table, image, and video tasks, within a unified framework, allowing knowledge transfer and improving performance across different forms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If separate pre-trained models are trained for different question-answering forms, then each model can be optimized for its specific form, but resource consumption and time costs increase significantly
Solution Approach 1:
The patent combines multiple question-answering tasks (text-based, knowledge-based, table-based, image-based, and video-based QA) into a single unified pre-training framework. The model simultaneously learns from diverse QA forms during pre-training, enabling it to handle different QA types without requiring separate models, thus reducing resource consumption and training time while maintaining performance.
Solution Approach 2:
The patent creates a universal pre-trained model that can perform multiple question-answering functions across different modalities. The model is designed with multi-functionality to handle text, knowledge, tables, images, and videos within a single framework, eliminating the need for separate specialized models and achieving both optimization and efficiency.
2Reliability
If separate pre-trained models are trained for different question-answering forms, then each model can achieve specialized performance, but the complexity of the system increases
Solution Approach 1:
The patent merges multiple specialized models into a single unified model that handles text-based, knowledge-based, table-based, image-based, and video-based question answering. This consolidation reduces system complexity by eliminating the need to manage multiple separate models while maintaining specialized performance through multi-task learning during pre-training.
3Reliability
If separate pre-trained models are trained for different question-answering forms, then each model can be optimized independently, but time costs and training resources are consumed excessively
Solution Approach 1:
The patent performs preliminary joint pre-training on diverse question-answering tasks before fine-tuning for specific applications. By pre-training the model simultaneously on multiple QA forms (text, knowledge, tables, images, videos) in advance, the model acquires generalized capabilities that reduce the time needed for subsequent specialized training and deployment.
Data Source
Figure 1
Figure 2~3
Figure 4
AI summary
The present disclosure provides a method and apparatus for acquiring a pre-trained model, an electronic device and a storage medium, and relates to the fields such as deep learning, natural language processing, knowledge graph and intelligent voice. The method may include: acquiring a pre-training task set composed of M pre-training tasks, M being a positive integer greater than 1, the pre-training tasks including: N question-answering tasks corresponding to different question-answering forms, N being a positive integer greater than 1 and less than or equal to M; and jointly pre-training the pre-trained model according to the M pre-training tasks. By use of the solutions of the present disclosure, resource consumption may be reduced, and time costs may be saved.