Training method and device of multi-modal model, equipment and storage medium
CN117171573BActive Publication Date: 2026-07-24BEIJING ZITIAO NETWORK TECH CO LTD
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ZITIAO NETWORK TECH CO LTD
- Filing Date
- 2023-09-21
- Publication Date
- 2026-07-24
Smart Images

Figure CN117171573B_ABST
Abstract
Embodiments of the present disclosure relate to a multi-modal model training method, device and equipment and a storage medium. The multi-modal model training method comprises: obtaining paired image samples and text samples corresponding to a target task; performing region feature extraction and text recognition on the image samples to generate image region features and image text features respectively; inputting the text samples, the image region features and the image text features into a multi-modal pre-training model to generate output data corresponding to the target task; wherein the output data comprises output images and / or output texts; and training the multi-modal pre-training model based on sample true values corresponding to the output data and the image samples. According to the embodiments of the present disclosure, more effective information is provided for the multi-modal pre-training model to obtain a text-image alignment function, the number of training samples required for model training is reduced, and the model training efficiency is improved.
Need to check novelty before this filing date? Find Prior Art
Citation Information
Patent Citations
Multi-modal model training and image recognition method and device, and electronic equipment
CN114239760A
Pre-training model training method and device, pre-training model application method and device, electronic equipment and medium
CN116186545A