基于信息抽取的多模态表征方法和装置

By extracting information and processing structured triples, structurally inconsistent text descriptions are generated, which improves the structural semantic discrimination ability of multimodal models, solves the problem of insufficient structured semantic relationship discrimination in existing technologies, and achieves stronger semantic understanding and discrimination capabilities.

CN122173896BActive Publication Date: 2026-07-17NAT UNIV OF DEFENSE TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NAT UNIV OF DEFENSE TECH
Filing Date
2026-05-08
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing multimodal representation methods lack the ability to distinguish structured semantic relationships, making it difficult to differentiate the structural relationships between subjects, predicates, and objects in text. Furthermore, they do not explicitly introduce structured knowledge as model input or constraints, resulting in insufficient understanding of structural semantics by the model.

Method used

By extracting structured semantic units from multimodal data through information extraction, constructing a set of structured triples and perturbing the structure and semantics, generating structurally inconsistent text descriptions, encoding image and text data respectively, fusing global semantic and structural semantic representations, and applying structure-aware discriminative constraints to improve the model's structural semantic discrimination ability.

Benefits of technology

It significantly improves the robustness and deep understanding of multimodal models for complex semantic structures. Through explicit structured information processing and structure-aware discrimination, it enhances the model's ability to distinguish between structurally consistent and inconsistent samples, thereby strengthening the model's structural semantic understanding capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122173896B_ABST
    Figure CN122173896B_ABST
Patent Text Reader

Abstract

本申请涉及一种基于信息抽取的多模态表征方法和装置。所述方法包括:获取图像和文本描述;对文本进行信息抽取,得到包含主体、关系、客体的结构化语义单元;构建规范化三元组集合;基于三元组进行结构语义扰动,生成结构不一致文本;分别编码图像、文本和三元组,得到图像表征、文本全局语义表征和结构语义表征;融合文本全局语义表征与结构语义表征,得到结构增强文本表征;计算图像与结构一致 / 不一致文本增强表征的匹配度,并执行结构感知判别约束,使结构一致样本匹配度更高。本方法通过显式引入结构化语义并构造高质量对比样本,有效提升了多模态表征对结构语义差异的区分能力。
Need to check novelty before this filing date? Find Prior Art