Camera lens defect detection method and device, storage medium and product

By combining multimodal imaging data and deep learning models, the problems of low efficiency and accuracy in industrial camera lens coating defect detection have been solved, and automated and accurate defect detection has been achieved.

CN120525863BActive Publication Date: 2025-10-17SHENZHEN TIANDING AUTOMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510955788.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-17
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

In the existing technology, industrial camera lens coating defect detection relies on manual visual inspection, which has the problems of low efficiency, low accuracy, and difficulty in detecting defects such as tiny scratches.

Method used

A multimodal imaging data acquisition method is adopted, combining RGB modality, polarized light modality and white light interferometer data. Defect detection is performed through a deep learning model, and feature fusion is performed using a multi-head self-attention mechanism and a cross-modal attention module to output multi-task detection results of defect area, type and parameters.

Benefits of technology

It realizes the automation, high efficiency and accuracy of industrial camera lens defect detection, can detect tiny scratches and other defects, replaces the traditional manual step-by-step inspection process, and improves detection efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525863B_ABST
    Figure CN120525863B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a camera lens defect detection method and device, a storage medium and a product, wherein the method comprises the following steps: obtaining an initial detection image and initial detection data of a lens to be detected; preprocessing the initial detection image and the initial detection data to obtain a basic detection image of the lens to be detected and basic detection data of the lens to be detected; inputting the basic detection image and the basic detection data into a multi-modal encoder of a defect detection model for encoding processing to obtain multi-modal fusion features corresponding to the lens to be detected; inputting the multi-modal fusion features into a multi-task decoder of the defect detection model for decoding processing to obtain multi-task detection results corresponding to the lens to be detected; by using the application, the automation of camera lens defect detection can be realized, and the efficiency and accuracy of camera lens defect detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, and in particular to a camera lens defect detection method and device, a storage medium and a product. BACKGROUND

[0002] Industrial cameras are image acquisition devices specially designed for industrial applications, with characteristics such as high reliability, high precision, and strong environmental adaptability, and are widely used in fields such as automated detection, machine vision, and intelligent manufacturing. Among them, industrial camera lens coating is a technology that uses optical interference principles to coat one or more thin films on the surface of the lens to improve the optical performance of the lens. Due to oversight in production and processing, or scratches caused by collisions and friction during transportation, it is easy to cause defects in the lens coating, which seriously affects product quality. Therefore, how to efficiently detect defects in industrial camera lenses is a technical problem worthy of study. SUMMARY

[0003] The embodiments of the present application provide a camera lens defect detection method and device, a storage medium and a product, which can realize the automation of camera lens defect detection and improve the efficiency and accuracy of camera lens defect detection.

[0004] In a first aspect, the embodiments of the present application provide a camera lens defect detection method, which comprises: obtaining an initial detection image and initial detection data of a lens to be detected; the initial detection image comprises an image of an RGB modality of the lens to be detected and an image of a polarized light modality of the lens to be detected; the initial detection data comprises data obtained by measuring the lens to be detected using a white light interferometer; preprocessing the initial detection image and the initial detection data to obtain a basic detection image of the lens to be detected and basic detection data of the lens to be detected; geometrically aligning the basic detection image and the basic detection data; inputting the basic detection image and the basic detection data into a multi-modal encoder of a defect detection model for encoding processing to obtain multi-modal fusion features corresponding to the lens to be detected; the multi-modal encoder comprises a window attention module based on a shift window and a cross-modal attention module based on a multi-head self-attention mechanism; inputting the multi-modal fusion features into a multi-task decoder of the defect detection model for decoding processing to obtain a multi-task detection result corresponding to the lens to be detected; the multi-task detection result comprises a first detection result, a second detection result, and a third detection result; the first detection result is used to indicate a defect region of the lens to be detected, the second detection result is used to indicate a defect type of the lens to be detected, and the third detection result is used to indicate a defect parameter of the lens to be detected.

[0005] In an implementation manner, the specific implementation manner that the initial detection image and the initial detection data are preprocessed to obtain the base detection image corresponding to the lens to be detected and the base detection data can be: the initial detection image and the initial detection data are subjected to feature matching processing to obtain an intermediate detection image of the lens to be detected and intermediate detection data of the lens to be detected; the feature matching processing includes at least one of scale-invariant feature transformation processing and thin plate spline transformation processing; the intermediate detection image and the intermediate detection data are down-sampled through two-level convolution operations to obtain the base detection image of the lens to be detected and the base detection data of the lens to be detected.

[0006] In an implementation manner, the specific implementation manner that the base detection image and the base detection data are input into the multi-modal encoder of the defect detection model for encoding processing to obtain the multi-modal fusion feature corresponding to the lens to be detected can be: the base detection image and the base detection data are divided into a plurality of non-overlapping local windows by using a window attention module based on a shift window, and a self-attention feature corresponding to each local window is determined; the self-attention features corresponding to the plurality of local windows are fused by using a cross-modal attention module based on a multi-head self-attention mechanism to obtain the multi-modal fusion feature corresponding to the lens to be detected.

[0007] In an implementation manner, the cross-modal attention module based on the multi-head self-attention mechanism includes a query component, a key component and a value component; the query component corresponds to the RGB modality, the key component corresponds to the polarized light modality, and the value component corresponds to the interference imaging data.

[0008] In an implementation manner, the method can further include: based on the training sample, the initial detection model is iteratively trained by the masking autoencoder until iteration termination, to obtain the base detection model; the training sample is obtained by masking the unlabeled data; the base detection model is fine-tuned based on a hybrid loss function to obtain the defect detection model; the hybrid loss function includes a first loss function for a defect area, a second loss function for a defect type and a third loss function for a defect parameter.

[0009] In an implementation manner, the first loss function includes a Dice loss function and a focal loss function; the second loss function includes a label smoothing cross-entropy loss function; and the third loss function includes a smooth L1 loss function.

[0010] In a second aspect, an embodiment of the present application provides a camera lens defect detection device, the device comprising: an acquisition unit configured to acquire an initial detection image and initial detection data of a lens to be detected; the initial detection image comprising an image in an RGB modality of the lens to be detected and an image in a polarized light modality of the lens to be detected; the initial detection data comprising data obtained by measuring the lens to be detected using a white light interferometer; a preprocessing unit configured to preprocess the initial detection image and the initial detection data to obtain a basic detection image of the lens to be detected and basic detection data of the lens to be detected; the basic detection image being geometrically aligned with the basic detection data; a feature fusion unit configured to input the basic detection image and the basic detection data into a multi-modal encoder of a defect detection model for encoding processing to obtain multi-modal fusion features corresponding to the lens to be detected; the multi-modal encoder comprising a window attention module based on a shift window and a cross-modal attention module based on a multi-head self-attention mechanism; a defect detection unit configured to input the multi-modal fusion features into a multi-task decoder of the defect detection model for decoding processing to obtain multi-task detection results corresponding to the lens to be detected; the multi-task detection results comprising a first detection result, a second detection result, and a third detection result; the first detection result being used to indicate a defect region of the lens to be detected, the second detection result being used to indicate a defect type of the lens to be detected, and the third detection result being used to indicate a defect parameter of the lens to be detected.

[0011] In a third aspect, an embodiment of the present application provides another camera lens defect detection device, the camera lens defect detection device comprising a memory and a processor, the memory being configured to store a computer program, and the processor being configured to run the computer program to perform the method of the first aspect.

[0012] In a fourth aspect, an embodiment of the present application provides a computer storage medium storing a computer program, and when the computer program is run by a processor, the method of the first aspect is implemented.

[0013] In a fifth aspect, an embodiment of the present application provides a computer program product comprising a computer program, and when the computer program is run by a processor, the method of the first aspect is implemented.

[0014] By implementing the embodiments of the present application, on the one hand, the data in different modalities of the lens to be detected can be fully utilized to realize deep fusion of the physical features and the image features of the lens to be detected, which is conducive to improving the accuracy and robustness of camera lens defect detection; on the other hand, the detection results of scratch segmentation, defect type, and defect geometric parameters (such as length, width, and depth) can be output synchronously to replace the traditional manual step-by-step detection process, which is conducive to improving the efficiency of camera lens defect detection. BRIEF DESCRIPTION OF DRAWINGS

[0015] Figure 1A system architecture schematic diagram for application in an embodiment of the present application;

[0016] Figure 2 A flowchart of a camera lens defect detection method provided in an embodiment of the present application;

[0017] Figure 3 A flowchart of another camera lens defect detection method provided in an embodiment of the present application;

[0018] Figure 4 A structural schematic diagram of a camera lens defect detection device provided in an embodiment of the present application;

[0019] Figure 5 A structural schematic diagram of another camera lens defect detection device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.

[0021] Under the background of increasing demand for industrial camera applications, the quality requirements for industrial camera lenses are continuously improved. Among them, the coating of industrial camera lenses is a key technology to improve optical performance and adapt to industrial environments, which directly affects the imaging clarity, durability and environmental adaptability. Due to possible omissions in production and processing links and inevitable collisions and friction during transportation, scratches and other defects are often left on the surface of the lens coating, which in turn affects the overall quality of the product. Therefore, detecting lens coating defects is the key to ensuring product quality.

[0022] Currently, the detection of industrial camera lens coating defects largely relies on manual visual inspection. Although this method can find obvious defects to some extent, it has many limitations, such as missing small scratches and being unable to detect micron or sub-micron scratches. The accuracy of manual visual inspection is highly dependent on the experience, vision and attention of the inspector, and is easily affected by subjective factors, leading to inconsistency and unreliability of the detection results. In addition, manual inspection is low in efficiency and difficult to meet the needs of large-scale production and high-precision detection.

[0023] Therefore, the embodiments of the present application provide a camera lens defect detection method, device and computer storage medium, which can realize the automatic detection of industrial camera lens defects and improve the efficiency and accuracy of industrial camera lens defect detection.

[0024] Referring to Figure 1 , Figure 1 is a schematic diagram of a system architecture applied to an embodiment of the present application. As shown in Figure 1 , the system architecture includes a computer device 101 and a detection device 102.

[0025] Optionally, the computer device 101 can be a terminal device or a server, and the computer device 101 can establish a connection with the detection device 102. The detection device 102 can be configured to acquire multi-modal imaging data of a lens to be detected. Optionally, the detection device 102 can be configured with a shooting device configured to acquire an RGB image of the lens to be detected, a polarized light imaging device configured to acquire a polarized light image of the lens to be detected, and a white light interferometer configured to acquire a plurality of topographic data of the lens to be detected. The polarized light imaging device can include a rotating polarizer and a linear polarization camera. The linear polarization camera can be equipped with a plurality of polarization direction filters (such as 0°, 45°, 90°, and 135°), which can capture linearly polarized light of different directions to highlight scratches on the lens coating surface. The white light interferometer can measure the nanoscale surface topography of the lens to be detected to accurately detect the microscopic topography of the lens surface.

[0026] The multi-modal imaging data of the lens to be detected acquired by the detection device 102 can be sent to the computer device 101, so that the computer device 101 performs defect detection on the lens to be detected based on the received data.

[0027] Referring to Figure 2 , Figure 2 is a flowchart of a camera lens defect detection method provided by an embodiment of the present application. As shown in Figure 2 , the method can be executed by a computer device, and the method can include but is not limited to the following steps:

[0028] S201, acquiring an initial detection image and initial detection data of a lens to be detected.

[0029] The initial detection image corresponding to the lens to be detected can include an image of the lens to be detected in an RGB modality and an image of the lens to be detected in a polarized light modality, and the initial detection data can include data obtained by measuring the lens to be detected by a white light interferometer. The image in the polarized light modality can include multi-angle polarized light images, such as 0°, 45°, 90°, and 135° polarized light images. Optionally, the initial detection image can also include an image of the lens to be detected in an infrared modality. The image in the infrared modality can be an image obtained by using near-infrared imaging technology. Here, the image in the polarized light modality can highlight scratches on the surface of the lens to be detected, the data obtained by the white light interferometer can be used to detect the depth of the scratches, and the image in the infrared modality can be used to detect subsurface defects (such as bubbles).

[0030] In an implementation manner, the initial detection image and the initial detection data corresponding to the lens to be detected can be sent by a detection device to a computer device. The detection device can be designed with a six-degree-of-freedom adjusting platform mechanical structure, compatible with planar and curved lenses, capable of automatic focusing and multi-angle imaging, and can also use a shockproof optical platform to reduce the influence of environmental vibration on interference measurement, so as to obtain more accurate detection data.

[0031] S202, pre-processing the initial detection image and the initial detection data to obtain a basic detection image and basic detection data corresponding to the lens to be detected.

[0032] In the embodiments of the present application, the basic detection image can be understood as an image obtained after pre-processing the initial detection image, and the basic detection data can be understood as data obtained after pre-processing the initial detection data. The basic detection image and the basic detection data are geometrically aligned. That is, the pre-processed images of different modalities of the lens to be detected correspond to the pre-processed measurement data of the lens to be detected in space, and can be processed in the same coordinate system. In this way, the detection images and detection data obtained under different conditions can be effectively integrated and analyzed.

[0033] In an implementation manner, the pre-processing of the initial detection image and the initial detection data can be feature matching processing of the initial detection image and the initial detection data to obtain an intermediate detection image of the lens to be detected and intermediate detection data of the lens to be detected, and then performing two-level convolution operation on the intermediate detection image and the intermediate detection data to down-sample to obtain the basic detection image and the basic detection data of the lens to be detected. Optionally, the feature matching processing of the initial detection image and the initial detection data can be at least one of scale-invariant feature transform (SIFT) processing and thin plate spline (TPS) processing. The SIFT processing can accurately detect and describe the key points in the image, and the TPS processing can correct local deformation. By SIFT processing to realize feature point matching, corresponding point pairs between different images can be found, and then a TPS transformation model is used for non-rigid registration of the images, so that the displacement difference between images and data of different modalities can be eliminated.

[0034] For example, taking SIFT processing as an example, the process of feature matching processing of the initial detection image and the initial detection data can be represented by the following formula:

[0035] (1)

[0036] wherein, , may represent the initial detection image of the RGB modality, may represent the number of pixels in the vertical direction, may represent the number of pixels in the horizontal direction, 3 may represent three color channels, may represent the intermediate detection image of the RGB modality. , may represent the initial detection image of the polarized light modality, may represent the polarization angle, which may have a value of, for example, , may represent the intermediate detection image of the polarized light modality. , may represent the initial detection data of the interferometric imaging, may represent the vertical resolution of the single-layer interferometric imaging data (which may have a value of, for example, 1024), may represent the horizontal resolution of the single-layer interferometric imaging data (which may have a value of, for example, 1024), may represent the number of scanning layers in the depth direction, may represent the intermediate detection data of the interferometric imaging.

[0037] In the embodiments of the present application, after obtaining the intermediate detection images and the intermediate detection data of different modalities, the intermediate detection images and the intermediate detection data can be input into a ConvStem composed of two levels of convolution and layer normalization (layernorm) for processing. Specifically, for the input intermediate detection images and the intermediate detection data, a 7x7 convolution kernel can be used for convolution operation first, which can capture local features in the input data. After the convolution operation, layer normalization can be performed, which can normalize all activations of each sample, which is conducive to improving the stability of the model. After layer normalization, a Gaussian Error Linear Unit (GELU) can be applied, which is a nonlinear activation function that can help the model learn more complex features. Further, after the activation function, a 3x3 convolution kernel can be used for the second convolution operation, which helps to further extract and combine features. Finally, layer normalization processing is performed again to ensure the stability of the data. Through two convolution operations and two layer normalization processes, the input data can be effectively pre-processed and features can be extracted, so as to provide more meaningful feature representation for subsequent model processing.

[0038] In some embodiments, the intermediate detection images and the intermediate detection data can be divided into small blocks after being divided, which can be fixed size (e.g., 16x16), and can be referred to as image blocks or data blocks, as inputs of the ConvStem consisting of two levels of convolution and layer normalization. Each input block can be embedded into a low-dimensional (e.g., 128-dimensional) feature vector by the ConvStem consisting of two levels of convolution and layer normalization, which can be used for subsequent visual tasks such as classification, detection, etc. In this way, the features of the multi-modal data can be effectively extracted while preserving the local spatial information and adapting to the high-resolution characteristics of industrial images.

[0039] Exemplarily, the process of inputting the intermediate detection images and the intermediate detection data into the ConvStem consisting of two levels of convolution and layer normalization for processing can be represented by the following formula:

[0040] For the RGB modality:

[0041] (2)

[0042] wherein, may represent the intermediate detection image of the RGB modality, may represent the first result obtained after the first level of convolution layer processing of the intermediate detection image of the RGB modality, may represent the second result obtained by layer normalization on the first result, may represent the third result obtained by applying the GELU nonlinear activation function to the second result, may represent the fourth result obtained by the second convolution operation on the third result, may represent the basic detection image of the RGB modality obtained after layer normalization on the fourth result.

[0043] For the polarized light modality:

[0044] (3)

[0045] wherein, may represent the intermediate detection image of the polarized light modality, may represent the fifth result obtained after the first level of convolution layer processing of the intermediate detection image of the polarized light modality, may represent the sixth result obtained by layer normalization on the fifth result, may represent the seventh result obtained by applying the GELU nonlinear activation function to the sixth result, may represent the eighth result obtained by the second convolution operation on the seventh result, may represent the basic detection image of the polarized light modality obtained after layer normalization on the eighth result.

[0046] For the interference imaging data:

[0047] (4)

[0048] wherein, may represent intermediate detection data of the interference imaging, may represent a ninth result obtained by processing the intermediate detection data of the interference imaging through a first convolutional layer, may represent a tenth result obtained by layer normalization on the ninth result, may represent an eleventh result obtained by applying a GELU nonlinear activation function to the tenth result, may represent a twelfth result obtained by a second convolutional operation on the eleventh result, may represent a basic detection image of the interference imaging obtained by layer normalization on the twelfth result.

[0049] S203, input the basic detection image and the basic detection data into the multi-modal encoder of the defect detection model for encoding processing to obtain multi-modal fusion features corresponding to the lens to be detected.

[0050] Wherein, the defect detection model can be a model with the characteristics of the Transformer architecture, and the multi-modal encoder of the defect detection model can include a shift window-based attention module (swin transformer block) and a cross-modal attention module based on a multi-head self-attention mechanism.

[0051] In some embodiments, before inputting the basic detection image and the basic detection data into the multi-modal encoder of the defect detection model, the basic detection image and the basic detection data can be spliced to obtain spliced data. For example, the splicing process of the basic detection image and the basic detection data can be represented by the following formula:

[0052] (5)

[0053] wherein, may represent spliced data, may represent a basic detection image of the RGB modality, may represent a basic detection image of the polarized light modality, may represent a basic detection image of the interference imaging.

[0054] In an implementation, the multi-modal encoder can combine hierarchical down-sampling (e.g., 4x, 8x, 16x) to extract multi-scale features in stages. For each stage, a window attention module based on a shifted window mechanism can be utilized to divide the base detection image and the base detection data into multiple non-overlapping local windows, and then determine the self-attention features corresponding to each local window. Here, the shifted window mechanism is introduced to further realize the fusion of cross-window information. A cross-modal attention module based on a multi-head self-attention mechanism can be used to fuse the self-attention features corresponding to multiple local windows to obtain multi-modal fusion features of the detection shot. The cross-modal attention module based on the multi-head self-attention mechanism can simultaneously process and integrate information from different modalities (e.g., images of the RGB modality, images of the polarized light modality, and interference imaging data). Specifically, the cross-modal attention module based on the multi-head self-attention mechanism can include three components: query, key, and value. The query can come from the image of the RGB modality, and the key and value can come from the image of the polarized light modality and the interference imaging data. The similarity or matching degree between the query vector k and the key vector can be calculated using the query vector k to generate attention weights. These weights can represent the relevance of each value vector to the current query. The weighted sum of the value vectors according to the attention weights can obtain a new feature representation that integrates information from multiple modalities. In addition, the cross-modal attention module based on the multi-head self-attention mechanism can combine hierarchical skip connections in deep networks. Through the hierarchical skip connections, the model can retain the detailed information from the early layers, which helps to improve the detection performance of the defect detection model, and finally enables the model to output multi-modal feature representations with global semantics and local precision.

[0055] In an implementation, the process of inputting the base detection image and the base detection data into the multi-modal encoder of the defect detection model for encoding processing can be summarized as the following steps:

[0056] Step 1: Perform layer normalization on the input base detection image and base detection data. The input to the multi-modal encoder of the defect detection model can be the spliced data of the base detection image and the base detection data . The normalization processing of the spliced data can be represented by the following formula:

[0057] LayerNorm (6)

[0058] Step 2: Calculate the self-attention features within the local window through the window-based multi-head self-attention (W-MSA) mechanism. The input spliced data can be divided into multiple non-overlapping local windows, and each window can contain a fixed number of elements. For each window, the query matrix , the key value matrix and the value matrix can be generated through linear transformation:

[0059] (7)

[0060] (8)

[0061] (9)

[0062] wherein , and can represent the linear transformation matrix, can be a linear transformation matrix associated with the image information of the RGB modality, can be a linear transformation matrix associated with the image information of the polarized light modality, can be a linear transformation matrix associated with the interference imaging data, can represent the normalized spliced data.

[0063] The self-attention features can be represented by the following formula:

[0064] (10)

[0065] wherein can represent the column number of the query matrix and the key value matrix , i.e. the vector dimension, can represent the learnable relative position bias, which can be used to capture the relative position relationship between elements.

[0066] Step 3: Perform residual connection between the self-attention features and the input spliced data , and then perform layer normalization, which can be represented by the following formula:

[0067] (11)

[0068] Step 4: Perform multi-layer perceptron (MLP) processing on to further extract and fuse feature information:

[0069] (12)

[0070] wherein, may represent an activation function, may represent a weight matrix, may represent a bias vector.

[0071] Step 5: Finally, a residual connection is performed again, and the output of the multi-modal encoder of the defect detection model can be obtained:

[0072] (13)

[0073] In this way, rich local and global features can be captured while maintaining computational efficiency.

[0074] S204, input the multi-modal fusion feature into the multi-task decoder of the defect detection model for decoding processing to obtain a multi-task detection result corresponding to the to-be-detected shot.

[0075] In an embodiment of the present application, the multi-task detection result obtained by detecting the to-be-detected shot by using the defect detection model can include a first detection result, a second detection result, and a third detection result. The first detection result can be used to indicate a defect region of the to-be-detected shot, the second detection result can be used to indicate a defect type of the to-be-detected shot, and the third detection result can be used to indicate a defect parameter of the to-be-detected shot. Optionally, the defect region of the to-be-detected shot can be represented by a pixel-level segmentation mask, and each pixel can be assigned a class label to identify a scratch or a non-scratch. The defect type of the to-be-detected shot can include, but is not limited to, scratches, bubbles, and stains, etc. The defect parameter of the to-be-detected shot can include, but is not limited to, the length, width, and depth of a defect (such as a scratch), etc.

[0076] In an implementation manner, the multi-task decoder of the defect detection model can adopt a progressive upsampling structure in the style of U-Net, and the spatial resolution of the feature map can be restored layer by layer through transposed convolution. In addition, each level of the multi-task decoder of the defect detection model can be connected to the corresponding level of the multi-modal encoder of the defect detection model through a skip connection, which helps to retain the detailed information passed from the multi-modal encoder, thereby facilitating the improvement of the accuracy of defect detection. For the final output multi-task detection result, the first detection result can be a pixel-level segmentation mask generated by a 1x1 convolution operation, for example, it can be binary, such as each pixel can be classified as a scratch or a non-scratch; the second detection result can be obtained by classifying the defect type through a fully connected layer; and the third detection result can be the specific physical parameter of the defect predicted by the regression branch.

[0077] In some embodiments, the multi-task detection results can share the same multi-modal fusion feature, that is, different detection tasks can reuse the same feature representation, thereby facilitating the improvement of the efficiency and performance of the defect detection model. Optionally, the defect detection model can also optimize each detection task through dynamic weight balancing, for example, the weights of each task can be automatically adjusted according to the importance or difficulty of the detection task to achieve the best detection performance. By jointly optimizing multiple detection tasks, more accurate defect detection and quantitative analysis can be achieved.

[0078] In some embodiments, the defect detection model can be obtained by training an initial detection model in combination with self-supervised learning and multi-task learning. Specifically, in the model training phase, the initial detection model can be iteratively trained based on the training samples through a masked autoencoder (MAE) until the iteration is terminated, to obtain a basic detection model. The MAE can randomly mask a part of the input image, and then train the model to reconstruct the masked image block, so that the model can learn the intrinsic structure and features of the image without external labels. In the model fine-tuning phase, the basic detection model can be fine-tuned based on a hybrid loss function, and the defect detection model can be obtained. The hybrid loss function can jointly optimize multiple detection tasks, such as segmentation, classification, and regression tasks. For the segmentation task, the Dice loss function and the focal loss function can be used. The Dice loss can be used to measure the degree of overlap between the predicted segmentation mask and the real mask, which helps to solve the problem of class imbalance. The focal loss can increase the weight of difficult-to-classify samples to improve the model's ability to recognize minority classes. For the classification task, the label smoothing cross entropy loss function can be used. Smoothing the label distribution can reduce the overfitting of the model, thereby facilitating the improvement of the generalization of the model. For the regression task, the smooth L1 loss function can be used. The smooth L1 loss function can be used to accurately predict the set parameters of defects, and has good robustness to outliers.

[0079] Optionally, during the model fine-tuning process, the encoder parameters can be frozen through a small sample learning strategy, and only the decoder and task head can be fine-tuned, so as to achieve high-performance detection under limited labeled data. Here, freezing the encoder parameters can prevent overfitting, because the encoder has learned a general feature representation through self-supervised pre-training. Fine-tuning only the decoder and task head parameters can better use different detection tasks.

[0080] For ease of understanding, the camera lens defect detection method provided by the embodiments of the present application is described below in combination with a specific algorithm processing flow. Please refer toFigure 3 , Figure 3 Another flowchart of a camera lens defect detection method provided by an embodiment of the present application is shown in FIG. 6. As shown in FIG. 6, the method can include the following steps: Figure 3

[0081] ①Obtain the RGB modality image, the polarized light modality image and the interferometer imaging data of the lens to be detected. The RGB modality image and the polarized light modality image of the lens to be detected can also be described as the initial detection image of the lens to be detected, and the interferometer imaging data of the lens to be detected can also be described as the initial detection data of the lens to be detected.

[0082] ②Extract SITF features from the data of each modality to capture the key information of the data of each modality.

[0083] ③Embed the extracted features through a ConvStem structure, and the embedded features can be spliced together to form a multi-modal feature representation. The ConvStem structure can include a 7x7 convolutional layer, layer normalization, a GELU activation function, a 3x3 convolutional layer, and layer normalization. Steps ② and ③ can be understood as a process of preprocessing the multi-modal data of the lens to be detected.

[0084] ④An improved Swin Transformer model is used to extract multi-scale features in stages through local window attention and hierarchical down-sampling. Each stage includes layer normalization, window self-attention, residual connection, layer normalization, MLP, and residual connection. The window self-attention can calculate local feature interactions and introduce a shift window mechanism to realize cross-window information fusion. When calculating the window self-attention, multi-modal physical features can be dynamically fused through multi-head attention, such as a query matrix which can come from the data of the RGB modality, a key matrix which can come from the data of the polarized light modality, and a value matrix which can come from the interferometer imaging data. The improved Swin Transformer model can extract and fuse multi-modal data at a deep level, and then output a multi-modal feature representation that has both global semantics and local precision.

[0085] ⑤The output of the improved Swin Transformer model is passed through a fully connected layer to obtain the final multi-task detection result. For example, for the segmentation task, the fully connected layer can assign a class label to each pixel, such as a class label indicating whether it is a scratch; for the classification task, the fully connected layer can generate a prediction score for each defect type; and for the regression task, the fully connected layer can generate a predicted specific physical parameter of the defect.

[0086] ​In addition, to improve the efficiency of defect detection of the to-be-detected lens, the embodiments of the present application can optimize the real-time inference of the defect detection model from the aspects of model lightweight and hardware acceleration. The model lightweight can include: ①adopting knowledge distillation to compress the model size; ②adopting channel pruning to remove redundant attention heads and quantizing the model parameters to 8-bit integers; and ③adopting INT8 quantization to convert the model parameters from floating-point numbers to low-precision representations (such as 8-bit integers). In this way, the size and calculation amount of the model can be reduced, so that the calculation resources required by the defect detection model during inference are less, thereby accelerating the inference speed. The hardware acceleration can include: ①adopting a heterogeneous architecture of a graphics processing unit (GPU) and a neural processing unit (NPU) to process multi-modal data in parallel; and ②combining FPGA to accelerate the interference image preprocessing. Through the synergistic effect of model lightweight and hardware acceleration, the inference speed can be significantly improved while maintaining high-precision detection.

[0087] The camera lens defect detection method provided by the embodiments of the present application can integrate the RGB modal image, the polarized light modal image and the interference imaging data synchronously, realize deep fusion of physical features and image features, and thus improve the accuracy of camera lens defect detection. In addition, the camera lens defect detection method can output the detection results of scratch segmentation, defect type and defect geometric parameters (such as length, width and depth) synchronously, replace the traditional manual step-by-step detection process, and thus improve the efficiency of camera lens defect detection.

[0088] Referring to Figure 4 , Figure 4 is a structural schematic diagram of a camera lens defect detection device provided by the embodiments of the present application. As shown in Figure 4 , the camera lens defect detection device 40 can include:

[0089] The acquisition unit 401 is configured to acquire an initial detection image and initial detection data of a to-be-detected lens. The initial detection image includes an image of the to-be-detected lens in an RGB modal and an image of the to-be-detected lens in a polarized light modal. The initial detection data includes data obtained by measuring the to-be-detected lens by using a white light interferometer.

[0090] The preprocessing unit 402 is configured to pre-process the initial detection image and the initial detection data to obtain a basic detection image of the to-be-detected lens and basic detection data of the to-be-detected lens. The basic detection image and the basic detection data are geometrically aligned.

[0091] The feature fusion unit 403 is configured to perform encoding processing on the basis detection image and the basis detection data by inputting the basis detection image and the basis detection data into a multi-modal encoder of the defect detection model, to obtain multi-modal fusion features corresponding to the lens to be detected; the multi-modal encoder includes a window attention module based on a shift window and a cross-modal attention module based on a multi-head self-attention mechanism.

[0092] The defect detection unit 404 is configured to perform decoding processing on the multi-modal fusion features by inputting the multi-modal fusion features into a multi-task decoder of the defect detection model, to obtain multi-task detection results corresponding to the lens to be detected; the multi-task detection results include a first detection result, a second detection result and a third detection result; the first detection result is used to indicate a defect region of the lens to be detected, the second detection result is used to indicate a defect type of the lens to be detected, and the third detection result is used to indicate a defect parameter of the lens to be detected.

[0093] In an implementation manner, the preprocessing unit 402 is further configured to perform feature matching processing on the initial detection image and the initial detection data, to obtain an intermediate detection image of the lens to be detected and intermediate detection data of the lens to be detected; the feature matching processing includes at least one of scale-invariant feature transform processing and thin-plate spline transform processing; the intermediate detection image and the intermediate detection data are down-sampled by two-level convolution operations, to obtain the basis detection image of the lens to be detected and basis detection data of the lens to be detected.

[0094] In an implementation manner, the feature fusion unit 403 is further configured to divide the basis detection image and the basis detection data into a plurality of non-overlapping local windows by using the window attention module based on the shift window, and determine self-attention features corresponding to each local window; the self-attention features corresponding to the plurality of local windows are fused by using the cross-modal attention module based on the multi-head self-attention mechanism, to obtain the multi-modal fusion features corresponding to the lens to be detected.

[0095] In an implementation manner, the cross-modal attention module based on the multi-head self-attention mechanism includes a query component, a key component and a value component; the query component corresponds to an RGB modality, the key component corresponds to a polarized light modality, and the value component corresponds to interference imaging data.

[0096] In an implementation manner, the camera lens defect detection apparatus can further include a pre-training unit configured to perform iterative training on an initial detection model by using a masking autoencoder based on training samples until iteration termination, to obtain a basis detection model; the training samples are obtained by masking unlabeled data; the basis detection model is fine-tuned based on a hybrid loss function, to obtain the defect detection model; the hybrid loss function includes a first loss function for a defect region, a second loss function for a defect type and a third loss function for a defect parameter.

[0097] It should be noted that, Figure 4The contents not mentioned in the corresponding embodiments and the specific implementation manners of the various steps can be referred to Figure 2 and Figure 3 the embodiments shown in the foregoing and will not be described here again.

[0098] Please refer to Figure 5 , Figure 5 is a structural schematic diagram of another camera lens defect detection device provided in an embodiment of the present application, specifically, as shown in Figure 5 , the camera lens defect detection device 50 can include a receiver 501, a transmitter 502, a memory 503 and a processor 504, and the receiver 501, the transmitter 502, the memory 503 and the processor 504 are connected through one or more communication buses.

[0099] The receiver 501 can be configured to receive data, for example, the receiver 501 can be configured to receive initial detection images and initial detection data of a lens to be detected. The transmitter 502 can be configured to transmit data.

[0100] The memory 503 can include read-only memory and random access memory, and provide instructions and data to the processor 504. A portion of the memory 503 can also include non-volatile random access memory. The processor 504 can be a central processing unit (CPU), and the processor 504 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), FPGAs or other programmable logic devices, discrete gates or transistor logic components, discrete hardware components, etc. The general-purpose processor can be a microprocessor, and optionally, the processor 504 can also be any conventional processor or the like. Among them:

[0101] The memory 503 is configured to store program instructions.

[0102] The processor 504 is configured to call the program instructions stored in the memory 503, so as to:

[0103] The initial detection image and initial detection data of the lens to be detected are acquired; the initial detection image includes an image of an RGB mode of the lens to be detected and an image of a polarized light mode of the lens to be detected; the initial detection data includes data obtained by measuring the lens to be detected by using a white light interferometer; the initial detection image and the initial detection data are preprocessed to obtain a basic detection image of the lens to be detected and basic detection data of the lens to be detected; the basic detection image is geometrically aligned with the basic detection data; the basic detection image and the basic detection data are input into a multi-modal encoder of a defect detection model for encoding processing to obtain multi-modal fusion features corresponding to the lens to be detected; the multi-modal encoder includes a window attention module based on a shift window and a cross-modal attention module based on a multi-head self-attention mechanism; the multi-modal fusion features are input into a multi-task decoder of the defect detection model for decoding processing to obtain multi-task detection results corresponding to the lens to be detected; the multi-task detection results include a first detection result, a second detection result and a third detection result; the first detection result is used to indicate a defect region of the lens to be detected, the second detection result is used to indicate a defect type of the lens to be detected, and the third detection result is used to indicate a defect parameter of the lens to be detected.

[0104] In an implementation manner, the processor 504 can further be configured to perform feature matching processing on the initial detection image and the initial detection data to obtain an intermediate detection image of the lens to be detected and intermediate detection data of the lens to be detected; the feature matching processing includes at least one of scale-invariant feature transform processing and thin-plate spline transform processing; and the intermediate detection image and the intermediate detection data are down-sampled by two-level convolution operation to obtain the basic detection image of the lens to be detected and the basic detection data of the lens to be detected.

[0105] In an implementation manner, the processor 504 can further be configured to divide the basic detection image and the basic detection data into a plurality of non-overlapping local windows by using the window attention module based on the shift window, and determine self-attention features corresponding to each local window; and the self-attention features corresponding to the plurality of local windows are fused by using the cross-modal attention module based on the multi-head self-attention mechanism to obtain the multi-modal fusion features corresponding to the lens to be detected.

[0106] In an implementation manner, the cross-modal attention module based on the multi-head self-attention mechanism includes a query component, a key component and a value component; the query component corresponds to the RGB mode, the key component corresponds to the polarized light mode, and the value component corresponds to the interference imaging data.

[0107] In an implementation manner, the processor 504 can also be configured to perform iterative training on the initial detection model by the masking autoencoder based on the training samples until iteration termination, to obtain a basic detection model; the training samples are obtained by masking the unlabeled data; and the basic detection model is fine-tuned based on a hybrid loss function to obtain a defect detection model; the hybrid loss function includes a first loss function for a defect region, a second loss function for a defect type, and a third loss function for a defect parameter.

[0108] In an implementation manner, the first loss function includes a Dice loss function and a focal loss function; the second loss function includes a label smoothing cross-entropy loss function; and the third loss function includes a smooth L1 loss function.

[0109] It should be noted that, Figure 5 The contents not mentioned in the corresponding embodiments and the specific implementation manners of the steps can be referred to the embodiments shown in Figure 2 and Figure 3 and the foregoing contents, which will not be described here.

[0110] The embodiments of the present application also provide a computer readable storage medium, which stores a computer program. The computer program, when executed by a processor, causes the processor to perform the steps executed in the method embodiments shown in Figure 2 and Figure 3 .

[0111] The embodiments of the present application also provide a computer program product, which includes a computer program. The computer program, when executed by a processor, implements the steps executed in the method embodiments shown in Figure 2 and Figure 3 .

[0112] The steps in the embodiments of the present application can be adjusted in sequence, combined, and reduced according to actual needs.

[0113] The modules in the embodiments of the present application can be combined, divided, and reduced according to actual needs.

[0114] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by a computer program instructing related hardware. The program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like.

[0115] The above disclosure is only some embodiments of the present application, of course, cannot be limited by this, those skilled in the art can understand that the implementation of all or part of the above processes, and the equivalent changes made by the claims of the present application, still belong to the scope covered by the application.

Claims

1. A camera lens defect detection method, characterized in that: The method comprises: Acquire an initial test image and initial test data of the lens to be tested; the initial test image includes an image of the lens to be tested in RGB mode and an image of the lens to be tested in polarized light mode; the initial test data includes data obtained by measuring the lens to be tested using a white light interferometer; Preprocessing the initial detection image and the initial detection data to obtain a basic detection image of the lens to be detected and basic detection data of the lens to be detected; geometrically aligning the basic detection image with the basic detection data; Inputting the basic inspection image and the basic inspection data into a multimodal encoder of a defect detection model for encoding processing to obtain a multimodal fusion feature corresponding to the shot to be inspected; the multimodal encoder includes a window attention module based on a shifted window and a cross-modal attention module based on a multi-head self-attention mechanism; The multimodal fusion feature is input into the multi-task decoder of the defect detection model for decoding processing to obtain a multi-task detection result corresponding to the lens to be detected; the multi-task detection result includes a first detection result, a second detection result, and a third detection result; the first detection result is used to indicate the defect area of ​​the lens to be detected, the second detection result is used to indicate the defect type of the lens to be detected, and the third detection result is used to indicate the defect parameters of the lens to be detected.

2. The method according to claim 1, wherein The preprocessing of the initial detection image and the initial detection data to obtain a basic detection image and basic detection data corresponding to the lens to be detected includes: performing feature matching processing on the initial detection image and the initial detection data to obtain an intermediate detection image and intermediate detection data of the lens to be detected; the feature matching processing includes at least one of a scale-invariant feature transformation processing and a thin plate spline transformation processing; The intermediate detection image and the intermediate detection data are downsampled through a two-stage convolution operation to obtain a basic detection image of the to-be-detected shot and basic detection data of the to-be-detected shot.

3. The method according to claim 1 or 2, wherein: Inputting the basic detection image and the basic detection data into a multimodal encoder of a defect detection model for encoding processing to obtain a multimodal fusion feature corresponding to the lens to be detected includes: Dividing the basic detection image and the basic detection data into a plurality of non-overlapping local windows using the shifted window-based window attention module, and determining a self-attention feature corresponding to each local window; The cross-modal attention module based on the multi-head self-attention mechanism is used to fuse the self-attention features corresponding to multiple local windows to obtain the multimodal fusion features corresponding to the shot to be detected.

4. The method according to claim 3, wherein The cross-modal attention module based on the multi-head self-attention mechanism includes a query component, a key component and a value component; the query component corresponds to the RGB modality, the key component corresponds to the polarized light modality, and the value component corresponds to the interference imaging data.

5. The method according to claim 4, wherein The method further comprises: Iteratively training the initial detection model through a masked autoencoder based on training samples until the iteration is terminated to obtain a basic detection model; the training samples are obtained by masking unlabeled data; The basic detection model is fine-tuned based on a hybrid loss function to obtain a defect detection model; the hybrid loss function includes a first loss function for the defect area, a second loss function for the defect type, and a third loss function for the defect parameter.

6. The method according to claim 5, wherein The first loss function includes a Dice loss function and a focal loss function; the second loss function includes a label smoothed cross entropy loss function; and the third loss function includes a smoothed L1 loss function.

7. A camera lens defect detection device, characterized in that: The device comprises: an acquisition unit, configured to acquire an initial detection image and initial detection data corresponding to the lens to be detected; the initial detection image includes an image in RGB mode corresponding to the lens to be detected and an image in polarized light mode corresponding to the lens to be detected; the initial detection data includes data obtained by measuring the lens to be detected using a white light interferometer; a preprocessing unit, configured to preprocess the initial detection image and the initial detection data to obtain a basic detection image and basic detection data corresponding to the lens to be detected; and geometrically align the basic detection image with the basic detection data; a feature fusion unit, configured to input the basic inspection image and the basic inspection data into a multimodal encoder of a defect detection model for encoding processing to obtain a multimodal fusion feature corresponding to the shot to be inspected; the multimodal encoder includes a window attention module based on a shifted window and a cross-modal attention module based on a multi-head self-attention mechanism; A defect detection unit is configured to input the multimodal fusion features into the multi-task decoder of the defect detection model for decoding processing to obtain a multi-task detection result corresponding to the lens to be detected; the multi-task detection result includes a first detection result, a second detection result, and a third detection result; the first detection result is used to indicate a defect area of ​​the lens to be detected, the second detection result is used to indicate a defect type of the lens to be detected, and the third detection result is used to indicate a defect parameter of the lens to be detected.

8. A camera lens defect detection device, characterized in that: The device includes a memory and a processor, the memory is used to store a computer program, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 6.

9. A computer storage medium, characterized in that The computer storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Method and system for analyzing defects in wafer manufacturing based on big data

    CN119580022A

  • Optical device defect detection method and system

    CN120125580A