一种基于多模态特征的文本关联检测方法及装置

By combining a multimodal feature collaboration mechanism and a semantic decoupling mechanism, and analyzing text images using visual, geometric, and semantic features, the robustness problem of text detection in complex scenarios is solved, and efficient and accurate text association detection is achieved.

CN121708605BActive Publication Date: 2026-07-17BEIJING INFORMATION SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing text detection methods lack robustness in complex scenarios, especially in occlusion, lighting variations, or multilingual scenarios. Bottom-up methods are computationally inefficient, while top-down methods perform poorly when dealing with irregular text. Furthermore, existing relational modeling techniques fail to fully explore the correlation between text semantics and visual features.

Method used

By analyzing text images through a collaborative mechanism based on visual, geometric, and semantic features, feature updates are performed using a multimodal feature collaboration mechanism and a semantic decoupling mechanism. Relationship decision calculations are combined to generate associated text boxes, thereby improving detection accuracy.

Benefits of technology

It improves text detection efficiency in complex layout scenarios, reduces semantic noise interference, preserves the clear semantic information of the text, and has the ability to correct errors, thereby improving the accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708605B_ABST
    Figure CN121708605B_ABST
Patent Text Reader

Abstract

本申请涉及一种基于多模态特征的文本关联检测方法、装置、电子设备及存储介质。该方法包括:对文本图像进行字符检测识别,得到初始多模态特征;根据多模态特征协同机制和语义解耦机制,对初始多模态特征进行更新,得到目标多模态特征;根据目标多模态特征进行关系决策计算,得到文本图像的关联文本框。本申请实施例的方法,可以基于视觉特征、几何特征和语义特征协同对文本图像中识别到的文本之间的关联关系进行分析,提高了排版复杂的文字场景下的文本检测效率。
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Large model-based OCR (Optical Character Recognition) method and system and storage medium

    CN118379742A

  • Pre-training for scene text detection

    US20240119743A1