Training method and device of multi-modal model, equipment and storage medium

CN117171573BActive Publication Date: 2026-07-24BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2023-09-21
Publication Date
2026-07-24

Smart Images

  • Figure CN117171573B_ABST
    Figure CN117171573B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to a multi-modal model training method, device and equipment and a storage medium. The multi-modal model training method comprises: obtaining paired image samples and text samples corresponding to a target task; performing region feature extraction and text recognition on the image samples to generate image region features and image text features respectively; inputting the text samples, the image region features and the image text features into a multi-modal pre-training model to generate output data corresponding to the target task; wherein the output data comprises output images and / or output texts; and training the multi-modal pre-training model based on sample true values corresponding to the output data and the image samples. According to the embodiments of the present disclosure, more effective information is provided for the multi-modal pre-training model to obtain a text-image alignment function, the number of training samples required for model training is reduced, and the model training efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Multi-modal model training and image recognition method and device, and electronic equipment

    CN114239760A

  • Pre-training model training method and device, pre-training model application method and device, electronic equipment and medium

    CN116186545A