Visual question answering model training method, answer generation method and related device

By obtaining sample instances and expanding instance sets to calculate the difference loss, the visual question answering model is adjusted to increase the difference in prediction results. This solves the problem of low accuracy in image answer predictions in existing models and improves the model's ability to distinguish question content.

CN116127024BActive Publication Date: 2025-09-12MASHANG CONSUMER FINANCE CO LTD +1
2 Cites 0 Cited by

Patent Information

Application Number
CN202211227692.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2025-09-12
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

Existing visual question answering models have low prediction accuracy for image-based answers, are easily affected by the superficial similarity of question types, and fail to effectively distinguish different question contents.

Method used

By obtaining sample instances and extended instance sets, the difference loss of the visual question answering model is calculated, and the model is adjusted to increase the difference between the prediction results of sample instances and extended instances, prompting the model to pay more attention to other information of the question and image and reduce confusion.

Benefits of technology

It improves the answer prediction accuracy of the visual question answering model, enhances the model's ability to distinguish superficially similar questions, and reduces confusion in prediction results.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The embodiments of this specification provide a training method, answer generation method, and related apparatus for a visual question-answering model. The method obtains a sample instance and an extended instance set corresponding to the sample instance; wherein the sample instance includes at least a sample image and a sample question; the extended instance set includes multiple extended instances; and the extended instances include at least an extended question and an extended image; the sample instance and the extended instances in the extended instance set are respectively input into a visual question-answering model to obtain corresponding prediction results; a difference loss is constructed based on the prediction results of the sample instance and the prediction results of the extended instance; the difference loss decreases as the difference between the prediction results of the sample instance and the prediction results of the extended instance increases; and the visual question-answering model is adjusted based on the difference loss. This improves the answer prediction accuracy of the visual question-answering model.
Need to check novelty before this filing date? Find Prior Art