基于工控机的视觉检测方法及系统

By combining the CLIP visual encoder and multi-head cross-attention mechanism with cross-modal semantic alignment and object matching, the accuracy degradation problem of unsupervised learning visual detection in new product lines and data privacy-restricted scenarios is solved, achieving high-precision detection of structural defects and logical anomalies, and adapting to the needs of flexible manufacturing scenarios.

CN121190425BActive Publication Date: 2026-07-17SHENZHEN SHENTONG LINGKONG INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN SHENTONG LINGKONG INTELLIGENT TECH CO LTD
Filing Date
2025-09-18
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing unsupervised learning autoencoder visual detection methods cannot obtain a sufficient number of normal samples in new product lines or data privacy-restricted scenarios, resulting in the model's inability to fully learn normal patterns, the failure of reconstruction error thresholds, a significant decrease in anomaly detection accuracy, and limited application scenarios.

Method used

The CLIP visual encoder is used to extract global and multi-scale local features of images to generate multi-granular text prompts. Deep interaction between text and visual features is achieved through multi-scale convolution kernels and multi-head cross-attention mechanism. Combined with cross-modal semantic alignment and object matching, structural defect and logical anomaly score maps are generated, and anomaly determination is performed through weighted fusion.

Benefits of technology

It enables high-precision detection of structural defects and logical anomalies without the need for target domain training data, adapts to flexible manufacturing scenarios, improves detection accuracy and real-time performance, and meets the needs of real-time quality inspection in industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190425B_ABST
    Figure CN121190425B_ABST
Patent Text Reader

Abstract

本申请涉及一种基于工控机的视觉检测方法及系统,其包括通过工业相机采集工业产品的原始图像,对原始图像进行预处理,得到标准化图像;基于标准化图像,通过CLIP视觉编码器提取图像的全局特征与多尺度局部特征,并生成多粒度文本提示,多粒度文本提示包括图像级全局提示和块级局部提示,通过多尺度卷积核与多头交叉注意力机制,实现文本和视觉特征的深度交互,获取像素级的结构缺陷异常得分图;基于标准化图像及预设的正常样本参照库,通过跨模态语义对齐与对象匹配生成逻辑异常得分图;采用加权融合对结构缺陷异常得分图与逻辑异常得分图进行融合评分,基于融合评分进行异常判定,基于异常判定在工控机界面进行显示。
Need to check novelty before this filing date? Find Prior Art