语义提取模型的训练方法及应用方法

By using a temperature-sparse attention mechanism and a hypergraph structure to complete missing modal features, and combining hybrid loss to optimize parameters, this approach solves the problems of low recognition accuracy and poor robustness of existing semantic extraction models in cross-modal data processing, and achieves efficient multimodal data representation and recognition.

CN121278298BActive Publication Date: 2026-07-17BEIJING ANBOTONG TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ANBOTONG TECH CO LTD
Filing Date
2025-10-09
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In existing semantic understanding solutions, semantic extraction models have poor semantic extraction accuracy, especially in cross-modal data processing where recognition accuracy is low. Furthermore, recognition performance drops significantly when data for a particular modality is missing. The single loss function leads to uneven distribution of embedding vectors and poor clustering results.

Method used

A temperature-sparse attention mechanism is used to calculate the attention weights between multimodal feature vectors, and hypergraph structures are used to complete the missing modal features when they are detected. The model parameters are optimized by combining uniformity loss and contrast loss. High-dimensional feature extraction and mapping are performed using EfficientNetV2-XL, DeBERTaV3-large and Audio Spectrogram Transformer models.

Benefits of technology

It improves the recognition accuracy and robustness of the semantic extraction model, maintains high performance even in the case of modality loss, enhances the fine-grained expression and clustering effect of multimodal data, and improves the recall and recognition accuracy of cross-modal retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121278298B_ABST
    Figure CN121278298B_ABST
Patent Text Reader

Abstract

本申请公开了一种语义提取模型的训练方法及应用方法,涉及人工智能技术领域。方法包括:获取多模态数据样本;通过预设编码器对多模态数据样本进行特征提取,并将提取后的特征映射至目标维度的语义空间,得到多模态数据样本对应的多模态特征向量;基于温度稀疏注意力机制,计算多模态特征向量之间的注意力权重;在检测到某一模态特征缺失的情况下,通过由多模态特征向量构建的超图结构,对缺失的模态特征进行补全,得到补全后的多模态特征向量;对补全后的多模态特征向量,计算由均匀性损失和对比损失构成的混合损失,并根据混合损失,对温度稀疏注意力机制中的参数进行优化。本申请提出了一种能够解决目前语义提取模型精度较差的训练方法。
Need to check novelty before this filing date? Find Prior Art