一种基于多模态特征的文本关联检测方法及装置
By combining a multimodal feature collaboration mechanism and a semantic decoupling mechanism, and analyzing text images using visual, geometric, and semantic features, the robustness problem of text detection in complex scenarios is solved, and efficient and accurate text association detection is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-07-17
AI Technical Summary
Existing text detection methods lack robustness in complex scenarios, especially in occlusion, lighting variations, or multilingual scenarios. Bottom-up methods are computationally inefficient, while top-down methods perform poorly when dealing with irregular text. Furthermore, existing relational modeling techniques fail to fully explore the correlation between text semantics and visual features.
By analyzing text images through a collaborative mechanism based on visual, geometric, and semantic features, feature updates are performed using a multimodal feature collaboration mechanism and a semantic decoupling mechanism. Relationship decision calculations are combined to generate associated text boxes, thereby improving detection accuracy.
It improves text detection efficiency in complex layout scenarios, reduces semantic noise interference, preserves the clear semantic information of the text, and has the ability to correct errors, thereby improving the accuracy of detection.
Smart Images

Figure CN121708605B_ABST
Abstract
Citation Information
Patent Citations
Large model-based OCR (Optical Character Recognition) method and system and storage medium
CN118379742A
Pre-training for scene text detection
US20240119743A1