语义提取模型的训练方法及应用方法
By using a temperature-sparse attention mechanism and a hypergraph structure to complete missing modal features, and combining hybrid loss to optimize parameters, this approach solves the problems of low recognition accuracy and poor robustness of existing semantic extraction models in cross-modal data processing, and achieves efficient multimodal data representation and recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING ANBOTONG TECH CO LTD
- Filing Date
- 2025-10-09
- Publication Date
- 2026-07-17
AI Technical Summary
In existing semantic understanding solutions, semantic extraction models have poor semantic extraction accuracy, especially in cross-modal data processing where recognition accuracy is low. Furthermore, recognition performance drops significantly when data for a particular modality is missing. The single loss function leads to uneven distribution of embedding vectors and poor clustering results.
A temperature-sparse attention mechanism is used to calculate the attention weights between multimodal feature vectors, and hypergraph structures are used to complete the missing modal features when they are detected. The model parameters are optimized by combining uniformity loss and contrast loss. High-dimensional feature extraction and mapping are performed using EfficientNetV2-XL, DeBERTaV3-large and Audio Spectrogram Transformer models.
It improves the recognition accuracy and robustness of the semantic extraction model, maintains high performance even in the case of modality loss, enhances the fine-grained expression and clustering effect of multimodal data, and improves the recall and recognition accuracy of cross-modal retrieval.
Smart Images

Figure CN121278298B_ABST