A cross-modal model training method based on variational guided image local information routing
CN122153462APending Publication Date: 2026-06-05GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU UNIVERSITY
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-05
Smart Images

Figure CN122153462A_ABST
Abstract
The application discloses a cross-modal model training method based on variational guided image local information routing and belongs to the technical field of image classification. The method comprises the following steps: in the first step, a combined zero sample learning data set is acquired and pretreated; in the second step, a cross-modal model based on variational guided image local information routing is constructed; in the third step, the model is trained and fitted, and the model is evaluated; and in the fourth step, the model is evaluated and deployed. Without changing the original training target and loss function, the visual word is discriminatively guided and sequentially injected into the cross-modal process, the denoised visual feature is added to the text feature in the multiple paths, the ideal embedding for the combination recognition is formed, and the generalization ability of the model in the combined zero sample task is improved.
Need to check novelty before this filing date? Find Prior Art