Flexible visual detection method for abnormal surface of complex product based on virtual-real comparison
Through the complex product surface abnormality detection method based on virtual and real comparison, the problems of high data dependence and poor adaptability to complex products in the existing technology are solved, and flexible and intelligent detection of surface abnormalities of complex products are achieved, improving the adaptability of detectors to new products.
Patent Information
- Application Number
- CN202510387814.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art has problems such as high data dependence, poor adaptability to complex products, and difficulty in dealing with ‘new categories’ products and abnormalities in the detection of surface abnormalities of complex products, which limits the further improvement of the level of inspection flexibility.
The flexible visual detection method of complex product surface surface abnormality based on virtual and real comparison is adopted. By determining the product to be inspected and its CAD model, shooting planning and template image acquisition, virtual and real alignment, shooting execution and image acquisition, virtual and real comparison data set construction, abnormal data set construction based on cut-paste method, virtual and real comparison detector training, and abnormal detection based on virtual and real comparison.
It realizes flexible and intelligent detection of surface abnormalities of complex products, alleviates the scarcity of abnormal data, enhances the detection ability of new types of abnormalities, and improves the adaptability of detectors to new products.
Smart Images

Figure CN120236140A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of visual inspection of product quality, and particularly relates to a flexible visual inspection method for surface anomalies of complex products based on virtual-real comparison. Background Art
[0002] In high-end equipment manufacturing fields such as aerospace and precision instruments, products usually have complex functions and shapes, and the production mode features variable product types and small batch sizes. To control product quality, an important task is to inspect surface anomalies such as defects and foreign objects, usually by means of machine vision technology. Due to factors such as the unpredictable occurrence location of anomalies and the complex shape of products, it is usually necessary to take images from multiple perspectives to complete the inspection. A typical flexible visual inspection system usually includes important components such as inspection planning, data acquisition, and data analysis. The camera is usually fixed at the end of an execution device (such as a robotic arm, etc.). After a given inspection task is assigned, viewpoint planning based on a digital model is usually carried out first, that is, to determine the set of camera shooting poses. During image acquisition, the generalized actuator is positioned in the product coordinate system, and the camera is moved to the specified pose according to the planned scheme for taking pictures. Finally, the acquired images are analyzed to achieve anomaly detection and pixel-level positioning.
[0003] Among them, image-based anomaly detection currently still mainly relies on two major categories of methods: anomaly representation learning and normal representation learning. Anomaly representation learning methods usually rely on a sufficient amount of high-quality labeled anomaly samples to obtain more controllable performance. Normal representation learning methods only learn from normal samples and determine whether a test sample is abnormal, and are currently a widely concerned category of methods.
[0004] However, these methods still face several important challenges in the detection of surface anomalies of complex products, which limit the further improvement of the inspection flexibility level. First, current anomaly detection methods can usually effectively detect products and anomalies seen in the training set, but the detection ability for new classes that have not been seen will be greatly reduced. However, in the variable product type and small batch manufacturing mode, new products have a relatively high occurrence frequency. At the same time, due to the extremely strong randomness of the occurrence time, location, and cause, etc., the surface anomaly features are extremely rich and difficult to predict. Second, the excellent performance of these data-driven methods strongly depends on a sufficient amount of high-quality and diverse samples and annotations. However, in fact, the acquisition and annotation of samples is a cumbersome and expensive process, which requires the aid of appropriate sensors, annotation tools, and experienced engineers. In addition, considering that the defective rate in a normal production process is usually very low, this results in a very small number of anomaly samples that can be obtained. Summary of the Invention
[0005] The object of the present invention is: in the process of visual inspection of surface anomalies of products with complex textures and shapes, in order to overcome the problems of high data dependence, poor adaptability to complex products, and difficulty in dealing with "new type" products and anomalies in the existing methods, and to further improve the flexibility and intelligence level of visual inspection, the present invention provides a flexible visual inspection method for surface anomalies of complex products based on virtual-real comparison.
[0006] The present invention is mainly realized through the following technical solutions: A flexible visual inspection method for surface anomalies of complex products based on virtual-real comparison, characterized by comprising the following steps:
[0007] Determine the product to be inspected and its CAD model. This method realizes the detection of surface anomalies of products based on virtual-real comparison, where "real" refers to the real product; "virtual" refers to the digital model corresponding to the product, usually the CAD model; the data form participating in the "comparison" is an image.
[0008] Shooting planning and template image acquisition. It is difficult to predict the position where anomalies occur on the product surface, so it is usually necessary to traverse all target surfaces to be inspected during its detection. For large and complex products, in order to completely obtain the target surface image, it is necessary to take images from multiple viewpoints. Since the number of eligible shooting schemes is huge, an efficient and high-quality scheme needs to be planned. This process is completed in the virtual model space. After determining the shooting viewpoints, virtual views can be generated through model rendering as template images.
[0009] Virtual-real alignment. Before performing shooting, in order to ensure the execution accuracy of the planned scheme, the alignment between the model space and the physical space should be maintained as much as possible. First, the relative pose relationship installed between the camera and the execution device (such as a robotic arm) should be calibrated and kept consistent with the device parameters in the virtual space. Second, based on the detection coordinate system (usually the product coordinate system), the pose information of the device should be accurately obtained. By using the camera to take a scene image, the pose of the camera and the execution device in the detection coordinate system can be measured or estimated.
[0010] Shooting execution and acquisition of images to be inspected. Shooting execution means controlling the execution device in the real physical space according to the planned scheme to collect images to be inspected viewpoint by viewpoint. On the premise of ensuring virtual-real alignment, the template image and the image to be inspected taken at the same viewpoint according to the shooting scheme are also aligned and can form a virtual-real image pair.
[0011] Construction of the virtual-real comparison dataset. A surface anomaly detection dataset based on virtual-real comparison (ADD-VRC) is made as the basis for detector training and evaluation. The ADD-VRC dataset includes multiple dimensions such as products, template domains, inspection domains, anomalies, and samples. In a sample with the organizational form of "template-inspection-annotation" in the dataset, it records which product it comes from, which data domains the template image and the inspection image are sampled from respectively, and which type of anomaly the sample contains. Different from the products with good texture consistency recorded in the MVTecAD dataset, this dataset focuses on products with large sizes and complex shapes, such as some assemblies containing multiple parts, which introduces more complex background interference to the anomaly detection task. During the product imaging process, various factors such as camera internal parameters, viewpoints, and ambient light will affect the information in the image, thus interfering with anomaly detection. For example, camera internal parameters and viewpoints affect the spatial distribution of geometric features in the image, the light source layout affects the spatial distribution of bright and dark areas in the image, and the light intensity affects the overall brightness of the image, etc. The images affected by the combined influence of different factors are regarded as coming from different data domains. In the setting for the virtual-real comparison method, both the template image domain and the inspection image domain need to be considered. At this time, the inspection image is obtained from real products, while the template image is rendered from the virtual space according to the set parameters such as camera internal parameters, viewpoints, and lighting. In terms of the anomaly dimension, it mainly includes normal classes and anomaly classes such as surface defects and excess materials. Considering the characteristics of extremely rich anomaly features and difficult acquisition of anomaly samples, an anomaly synthesis method is adopted. On the basis of collecting normal classes and image alignment, the algorithm is used to automatically batch synthesize anomaly samples and perform annotation.
[0012] Construct the ADD-VRC dataset based on the cut-paste method. Considering that the process of obtaining real samples is time-consuming and laborious, and may cause irreversible damage to the product, a virtual anomaly dataset is generated batch by batch, automatically, and parametrically based on image synthesis technology. Cut-paste is an image enhancement method that cuts a small image a from image A and pastes it into image B to simulate anomalies. For products with 3D CAD models, first, images are taken with a camera, and the camera poses (i.e., viewpoints) corresponding to the images are recorded through measurement or calibration methods. In the virtual space, parameters such as the material and texture of the model, camera intrinsics, viewpoint, and light source are set, and virtual images are rendered and paired and aligned with real images. For the above-aligned full-size images, cropping is performed according to the same window to obtain a larger number of aligned cropped images. Appropriate samples are selected from public datasets (such as MvTecAD) for foreground cropping and randomly pasted onto the cropped images to be inspected as surface defects or extraneous objects. During the cut-paste process, random scaling, rotation, and Poisson transformation are performed on the anomalies to achieve the continuity between the foreground and the background, and the richness of the anomaly size and shape, etc.
[0013] Training of the virtual-real comparison detector. Select an appropriate anomaly detector according to the task requirements. Based on the completed dataset creation, consider the detector characteristics for detector training. The trained detector will be used in the formal detection process.
[0014] Anomaly detection based on virtual-real comparison. Given a "virtual-real image" pair, where the virtual image is the template image and the real image is the image to be inspected, potential anomalies in the image to be inspected can be detected and pixel-level localization can be achieved. In the basic framework of anomaly detection based on virtual-real comparison, there are mainly two key modules: feature extraction and similarity evaluation. That is, first, domain-independent and anomaly-sensitive embedding features of the input image are calculated, and then the similarity between the template feature and the feature to be inspected is calculated element by element, and finally, the anomaly region in the image to be inspected can be located. The main body of this method is a deep neural network structure to achieve more adaptive anomaly detection. The feature extraction module is a deep neural network structure, and weight assignment can be achieved by using a pre-trained model or training from scratch. The similarity evaluation module can be a neural network structure or a classic distance calculation method (such as cosine similarity); when it is a neural network structure, the similarity evaluation function needs to be implemented by training from scratch on the target dataset.
[0015] Advantages of the present invention: The present invention is oriented towards the visual detection of anomalies on the surface of complex products, and can fully exploit the potential knowledge in the digital model of the products to alleviate the problem of scarce anomaly data in industrial detection, enhance the anomaly detection capabilities for known and new classes, and improve the adaptability of the detector to new products. The present invention can be applied to the visual detection scenario of anomalies on the surface of complex products with a production mode of multiple varieties and small batches, realizing a more flexible visual detection. Description of the Drawings
[0016] Figure 1 is the flowchart of the flexible visual detection method for anomalies on the surface of complex products based on virtual-real comparison of the present invention.
[0017] Figure 2 is the product and model diagram of the present invention.
[0018] Figure 3 is the flowchart of the acquisition of virtual-real image pairs and anomaly synthesis of the present invention.
[0019] Figure 4 is an example diagram of the dataset for anomaly detection on the surface of complex products of the present invention.
[0020] Figure 5 is the anomaly detection framework diagram based on virtual-real comparison of the present invention.
[0021] Figure 6 is the structural diagram of the virtual-real comparison method VRC-ImageBind of the present invention.
[0022] Figure 7 is the structural diagram of the virtual-real comparison method VRC-SNUNet of the present invention. Detailed Embodiment
[0023] The following further elaborates on the present invention in conjunction with the drawings and embodiments.
[0024] The implementation of the present invention includes the following steps ( Figure 1 ):
[0025] Determine the product to be inspected and its CAD model. The laboratory simulation verification piece used is as Figure 2 shown. Its bottom plate has a length and width of 1200 mm and 900 mm respectively, and the height difference of the assembly is about 400 mm. It takes the large flat bottom plate as the reference, consists of dozens of components and connectors with different shapes and sizes, has a complex product shape, and has a meter-level size, etc. Except for the bottom plate and the cable bracket, the rest of the components are all 3D printed. In the laboratory scenario, the verification piece is placed on a support table. After manually adjusting the material / texture and light source, the model view is rendered based on Blender to reduce the visual perception difference from the real scene.
[0026] Shooting Planning and Template Image Acquisition. Prepare the *.stp model file and *.obj model file of the product instance. In the planning system, import the *.stp model of the instance, and this model will be displayed in the visualization interface of the system. Since the shape of this simulation verification part instance is relatively complex, the surfaces to be detected are selected manually: click the left mouse button on the surface where detection is planned, and the surface information will be recorded and the recommended viewpoints will be automatically generated. After all the surfaces to be detected are selected, the visibility evaluation program can be called to generate the visibility matrix. All intermediate data will be uniformly saved in a directory named after the instance name and the current timestamp. For example, the viewpoint information will be mainly saved as a visibility.json file, which records the viewpoint-related information of this sampling, including camera internal parameters, model path, surface sampling point set, unit / dimension, viewpoint set, etc. Plan the viewpoints through the greedy search algorithm, that is, the set of camera shooting poses.
[0027] Virtual-Reality Alignment. Before performing shooting, to ensure the execution accuracy of the planned scheme, the alignment between the model space and the physical space should be maintained as much as possible. First, calibrate the relative pose relationship installed between the camera and the execution device (such as a robotic arm) and keep it consistent with the device parameters in the virtual space. Second, based on the detection coordinate system (usually the product coordinate system), accurately obtain the pose information of the device. In terms of image-based pose estimation, specifically, global images are collected from a series of viewpoints far from the product, and its field of view can cover the overall contour of the product and is centered on the specified prominent elements. Select two relatively large parts near the center of the product as prominent elements. Each part has 8 three-dimensional bounding box vertices, and their three-dimensional coordinates in the detection coordinate system are constant, but their pixel coordinates in the two-dimensional image projection are related to the viewpoints. First, detect the bounding box vertices of the prominent elements in the image, obtain their pixel coordinates, and calculate the pose of the camera in the detection coordinate system according to the camera internal parameters and the three-dimensional coordinates of the vertices in the detection coordinate system.
[0028] Shooting Execution and Image to be Inspected Acquisition. Shooting execution means controlling the execution device in the real physical space according to the planned scheme and collecting the images to be inspected viewpoint by viewpoint. On the premise of ensuring virtual-reality alignment, the template image and the image to be inspected taken at the same viewpoint according to the shooting scheme are also aligned and can form a virtual-reality image pair.
[0029] Construction of the real - virtual comparison dataset. Generate the ADD - VRC dataset as described above. It covers 2 types of products and 12 types of surface anomalies (including the normal type), and involves one real domain and one virtual domain. Take a complex assembly with multiple components in the laboratory as the experimental object, take full - size images of the assembly from multiple angles, and perform cropping after aligning the model space and the physical space. Extract the features (with a dimension of 1280) of the cropped images based on the pre - trained base model Imagebind_huge, and then use the K - means method to automatically cluster 110 images into 5 small categories and manually divide them into two large categories. Due to the large differences in features, these two large categories can be regarded as two different products, namely P1 and P2. As Figure 3 shown, synthesize anomalies for these two products respectively according to cut - paste. Specifically, select appropriate samples from the public datasets MVTec AD and ITODD as the raw materials for anomaly synthesis. Given the determined product and data domain, each type of anomaly contains 100 groups of samples. Among them, paste the objects in ITODD as extra objects on the normal product images. The synthesized anomaly samples are as Figure 4 shown. The template images of this dataset are sampled from the virtual domain, and the images to be inspected are sampled from the real domain. Take images from the real scene and regard them as the real domain R; obtain virtual images by rendering in the virtual space and regard them as the virtual domain V. For all products and domains, randomly and evenly divide the samples of each type of anomaly into two parts, namely sample set I and II. Sample set I will only be used for training, and sample set II will be used for verification or testing as needed. To study the performance of different methods on new types of anomalies, establish two anomaly sets, namely and
[0030]
[0031]
[0032] Training of Virtual-Reality Comparison Detector. Select a suitable anomaly detector according to the task requirements. Based on the completed dataset creation, consider the detector characteristics for detector training. The trained detector will be used in the formal detection process. Specifically, two virtual-reality comparison method detector instances, VRC-ImageBind and VRC-SNUNet, are implemented, representing the unsupervised and supervised modes respectively. The training parameters are set as follows: For VRC-SNUNet, each batch contains 16 groups of samples during training, the AdamW optimizer is used, and the initial learning rate is 0.001; iterate 100 rounds on the training set and perform validation every 10 rounds; use the parameters with the best performance on the validation set as the final model parameters. For VRC-ImageBind, use the Imagebind_huge pre-trained large model to extract image features and perform similarity evaluation to detect potential anomalies. This process is completely unsupervised and does not require training. The main parameters of the computer used are: Intel Core i9-10900K CPU, NVIDIA GeForce RTX 3090 GPU, the memory capacity is 32GB, and the operating system is Windows 10.
[0033] Anomaly Detection Based on Virtual-Reality Comparison. Given a "virtual-real image" pair, where the virtual image is the template image and the real image is the image to be inspected, potential anomalies in the image to be inspected can be detected and pixel-level localization can be achieved. As Figure 5 shown, in the basic framework of anomaly detection based on virtual-reality comparison, it mainly includes two key modules: feature extraction and similarity evaluation. That is, first calculate the domain-independent and anomaly-sensitive embedding features of the input image, and then calculate the similarity between the template feature and the feature to be inspected element by element, so as to finally locate the abnormal area in the image to be inspected. The main body of this method is a deep neural network structure to achieve more adaptive anomaly detection. The feature extraction module is a deep neural network structure, and the weight assignment can be achieved by using a pre-trained model or training from scratch. The similarity evaluation module can be a neural network structure or a classic distance calculation method (such as cosine similarity); when it is a neural network structure, the similarity evaluation function needs to be implemented by training from scratch on the target dataset. The details of the two detector instances, VRC-ImageBind and VRC-SNUNet, are as follows:
[0034] VRC-ImageBind is modified from AnomalyGPT. It only uses one visual feature extractor and initializes its parameters with the pre-trained Imagebind_huge model. Using the pre-trained model and a fixed similarity evaluation method, VRC-ImageBind runs in a completely unsupervised manner. As Figure 6As shown in the figure, during the inference process, the pre-trained feature extractor extracts features from the template image and the image to be tested. Then, by calculating the cosine similarity of these feature maps, an abnormal feature map is obtained. Through binary processing, an abnormal mask is then derived. VRC-ImageBind first calculates the abnormal feature map Map with the same size as the image to be tested and the value is 0 to 1, and then uses a threshold t to obtain the abnormal mask Mask, following the formula: Mask ij =1if Map ij ≥t else0. The pixel with a value of 1 in the anomaly mask indicates that the pixel is judged as an anomaly. It can be seen that the choice of threshold directly affects the accuracy of anomaly detection. The trained anomaly detector is inferred on the validation set and the IoU under different thresholds is counted. As the threshold increases from 0 to 1, the IoU shows a trend of increasing first and then decreasing, and reaches a peak somewhere in the middle. The threshold corresponding to the IoU peak is determined as the optimal threshold and used in subsequent testing and reasoning stages.
[0035] The change detection method SNUNet-CD is directly used for virtual-real comparison, called VRC-SNUNet, such as Figure 7 As shown in Figure 2. Unlike VRC-ImageBind, VRC-SNUNet requires training the feature extraction and similarity assessment modules from scratch. The feature extraction module is a Siamese network in which parameters are shared between the two branches. The template image and the image undergo separate multi-layer and multi-scale feature extraction, respectively, and the size of the feature map is gradually restored by upsampling.
[0036] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall fall within the protection scope of the present invention.
Claims
1. A flexible visual inspection method for abnormal surface of complex products based on virtual-real comparison, characterized in that: The method comprises the following steps: Step S100: Determine the product to be inspected and its CAD model. This method realizes product surface anomaly detection based on virtual-real comparison, where "real" refers to the real product; "virtual" refers to the digital model corresponding to the product, usually a CAD model; the data involved in the "comparison" is in the form of an image. Step S200: Shooting planning and template image acquisition. Plan a set of viewpoints for shooting the product surface in the virtual model space to cover all target surfaces to be inspected. After determining the shooting viewpoints, a virtual view can be generated by model rendering as a template image. Step S300: virtual-real alignment. Before shooting, the model space should be aligned with the physical space. First, the relative posture relationship between the camera and the execution device (such as a robotic arm) should be calibrated and kept consistent with the device parameters in the virtual space. Second, the posture information of the device should be accurately obtained based on the detection coordinate system (usually the product coordinate system). Step S400: Shooting execution and image acquisition to be inspected. According to the planned scheme, the execution device is controlled in the real physical space to collect the image to be inspected from each viewpoint. Under the premise of ensuring virtual-real alignment, the template image and the image to be inspected taken at the same viewpoint according to the shooting scheme are also aligned, and a virtual-real image pair can be formed. Step S500: Construction of virtual-real comparison dataset: Constructing a surface anomaly detection dataset for complex products as the basis for detector training and evaluation. Step S600: Training of the virtual-real comparison detector. Select a suitable anomaly detector according to the task requirements, and train the detector considering the detector characteristics based on the completed data set creation. The trained detector will be used in the formal detection process. Step S700: Anomaly detection based on virtual-real comparison. Given a "virtual-real image" pair, where the virtual image is the template image and the real image is the image to be detected, potential anomalies in the image to be detected can be detected and pixel-level positioning can be achieved.
2. The method for flexible visual inspection of complex product surface abnormalities based on virtual-real comparison according to claim 1 is characterized in that: In step S500, the constructed surface anomaly detection dataset for complex products involves multiple dimensions such as product, template domain, domain to be inspected, anomaly, sample, etc. In a sample organized as "template-to-be-inspected-annotated" in the dataset, it is recorded which product it comes from, which data domain its template image and image to be inspected are sampled from, and what type of anomaly is contained in the sample.
3. The method for flexible visual inspection of complex product surface abnormalities based on virtual-real comparison according to claim 1 is characterized in that: In step S500, the constructed surface anomaly detection dataset for complex products supports the evaluation of the performance of the detection method on new types of products and new types of anomalies. Here, new categories refer to categories that appear in the test set but not in the training set. For all products and domains, the samples of each type of anomaly are randomly and evenly divided into two parts, namely, sample sets I and II. Sample set I will only be used for training, and sample set II will be used for verification or testing as needed. Two independent anomaly sets are established, namely and 4. The method for flexible visual inspection of complex product surface abnormalities based on virtual-real comparison according to claim 1 is characterized in that: In step S500, a data synthesis method is used to construct a data set, including obtaining normal samples through image rendering and real shooting, obtaining abnormal foreground from a public anomaly detection data set, and realizing abnormal synthesis based on cut-paste technology.
5. The method for flexible visual inspection of complex product surface abnormalities based on virtual-real comparison according to claim 1 is characterized in that: In step S700, the basic framework of anomaly detection based on virtual-real comparison mainly includes two key modules: feature extraction and similarity evaluation. That is, the domain-independent and anomaly-sensitive embedded features of the input image are first calculated, and then the similarity between the template features and the features to be detected is calculated element by element, so as to finally locate the abnormal area in the image to be detected.
6. The basic framework of anomaly detection based on virtual-real comparison according to claim 4 is characterized in that: The feature extraction module is a deep neural network structure, and weight assignment can be achieved by using a pre-trained model or training from scratch. The similarity assessment module can be a neural network structure or a classical distance calculation method (such as cosine similarity); when it is a neural network structure, the similarity assessment function needs to be achieved by training from scratch on the target data set.