基于工控机的视觉检测方法及系统
By combining the CLIP visual encoder and multi-head cross-attention mechanism with cross-modal semantic alignment and object matching, the accuracy degradation problem of unsupervised learning visual detection in new product lines and data privacy-restricted scenarios is solved, achieving high-precision detection of structural defects and logical anomalies, and adapting to the needs of flexible manufacturing scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN SHENTONG LINGKONG INTELLIGENT TECH CO LTD
- Filing Date
- 2025-09-18
- Publication Date
- 2026-07-17
AI Technical Summary
Existing unsupervised learning autoencoder visual detection methods cannot obtain a sufficient number of normal samples in new product lines or data privacy-restricted scenarios, resulting in the model's inability to fully learn normal patterns, the failure of reconstruction error thresholds, a significant decrease in anomaly detection accuracy, and limited application scenarios.
The CLIP visual encoder is used to extract global and multi-scale local features of images to generate multi-granular text prompts. Deep interaction between text and visual features is achieved through multi-scale convolution kernels and multi-head cross-attention mechanism. Combined with cross-modal semantic alignment and object matching, structural defect and logical anomaly score maps are generated, and anomaly determination is performed through weighted fusion.
It enables high-precision detection of structural defects and logical anomalies without the need for target domain training data, adapts to flexible manufacturing scenarios, improves detection accuracy and real-time performance, and meets the needs of real-time quality inspection in industry.
Smart Images

Figure CN121190425B_ABST