基于信息抽取的多模态表征方法和装置
By extracting information and processing structured triples, structurally inconsistent text descriptions are generated, which improves the structural semantic discrimination ability of multimodal models, solves the problem of insufficient structured semantic relationship discrimination in existing technologies, and achieves stronger semantic understanding and discrimination capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2026-05-08
- Publication Date
- 2026-07-17
AI Technical Summary
Existing multimodal representation methods lack the ability to distinguish structured semantic relationships, making it difficult to differentiate the structural relationships between subjects, predicates, and objects in text. Furthermore, they do not explicitly introduce structured knowledge as model input or constraints, resulting in insufficient understanding of structural semantics by the model.
By extracting structured semantic units from multimodal data through information extraction, constructing a set of structured triples and perturbing the structure and semantics, generating structurally inconsistent text descriptions, encoding image and text data respectively, fusing global semantic and structural semantic representations, and applying structure-aware discriminative constraints to improve the model's structural semantic discrimination ability.
It significantly improves the robustness and deep understanding of multimodal models for complex semantic structures. Through explicit structured information processing and structure-aware discrimination, it enhances the model's ability to distinguish between structurally consistent and inconsistent samples, thereby strengthening the model's structural semantic understanding capabilities.
Smart Images

Figure CN122173896B_ABST