Lens defect detection method and system based on region segmentation and semantic guidance

Through the lens defect detection method based on area segmentation and semantic guidance, the problem of differences in regional background and texture characteristics in lens detection is solved, efficient and accurate defect detection is achieved, and detection efficiency and accuracy are improved.

CN120298359APending Publication Date: 2025-07-11SHENZHEN KAIPULE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510379286.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing lens defect detection methods are poor in the face of the regular differences in the background and texture characteristics of different regions due to different materials, and traditional methods have problems such as loss of category information and incomplete semantic information when detecting multi-category anomalies.

Method used

Detection methods based on region segmentation and semantic guidance are adopted, including region segmentation preprocessing, semantic network feature fusion, multi-scale feature reconstruction and classification model. High-resolution images are collected through industrial cameras, and Hough circular transformation partitioning, semantic guidance network and diffusion denoising model are used to detect defects in combination with multi-scale feature fusion and classification model.

Benefits of technology

Accurate detection of lens defects is realized, detection efficiency and accuracy are improved, lens geometric information and semantic characteristics are fully utilized, and the reconstruction ability and detection accuracy of defect areas are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298359A_ABST
    Figure CN120298359A_ABST
Patent Text Reader

Abstract

The invention discloses a lens defect detection method and system based on region segmentation and semantic guidance. The method comprises the following steps: preprocessing a lens defect image based on a region segmentation preprocessing method; feature fusion is carried out on the preprocessed image based on a semantic network and spatial perception feature fusion, a lens image is reconstructed, and features in a potential space are mapped back to a pixel space; performing feature extraction and similarity calculation on the reconstructed lens image to generate an abnormal defect distribution diagram, and performing defect positioning; the detected defect areas are classified through the classification model, lens defect information is output, and the defect information comprises defect positions, types and severity degrees. According to the method, the problem that different area backgrounds and texture features show regular differences due to different materials is solved, the geometric information and semantic features of the lens are fully utilized, accurate detection of the defect area is realized, and the reconstruction capability of the defect area is improved by adopting the semantic guidance network and multi-scale feature fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and industrial automation detection, and particularly to a lens defect detection method and system based on region segmentation and semantic guidance, which can be widely applied to fields such as optical manufacturing, electronic component production, and quality inspection. Background Art

[0002] Currently, lens defect detection mainly uses traditional image processing methods or detection algorithms based on deep learning.

[0003] Among them, traditional methods use edge detection, threshold segmentation, and feature extraction technologies to analyze lens surface defects. These methods are sensitive to light changes and the complex surface textures of lenses, and have poor robustness.

[0004] Deep learning methods, detection algorithms based on neural networks learn defect features through training data and have certain detection capabilities.

[0005] The above methods have the following defects:

[0006] 1. In an actual scenario, for the lens images captured by a camera, the backgrounds and texture features of different regions show regular differences due to different materials;

[0007] 2. Traditional anomaly detection methods are usually based on a single category and perform well in complex multi-category detection scenarios. However, when directly applied to multi-category anomaly detection, there are problems such as loss of category information and incomplete semantic information. Summary of the Invention

[0008] The main objective of the present invention is to propose a lens defect detection method and system based on region segmentation and semantic guidance, aiming to improve the efficiency and accuracy of lens defect detection.

[0009] To achieve the above objective, the present invention provides a lens defect detection method based on region segmentation and semantic guidance, and the method includes the following steps:

[0010] Step S10, preprocess the lens defect image based on a region segmentation preprocessing method;

[0011] Step S20, perform feature fusion on the preprocessed image based on semantic network and spatial perception feature fusion, reconstruct the lens image, and map the features in the latent space back to the pixel space;

[0012] Step S30, perform feature extraction and similarity calculation on the reconstructed lens image, generate an abnormal defect distribution map, and perform defect localization;

[0013] Step S40: Classify the detected defect areas using a classification model and output lens defect information, where the defect information includes the defect location, type, and severity.

[0014] A further technical solution of the present invention is that the step S10 includes:

[0015] Step S101: Collect high-resolution defect images of the lens through an industrial camera to ensure clear presentation of the detailed textures of different areas;

[0016] Step S102: Use the region segmentation method to perform image cutting and rectangular padding of the lens:

[0017] Adopt the Hough circle transform to detect the center coordinates of the smallest circle in the image;

[0018] Cut the image according to the center coordinates and set the inner and outer diameters of multiple rings for partitioning;

[0019] Cut the ring or circle and return the smallest circumscribed rectangle;

[0020] Step S103: Perform normalization and scaling processing on the lens image to adjust the image size. Among them, the normalization formula is:

[0021]

[0022] Where I(x, y) is the pixel value of the input image; μ is the mean of all pixels, used to standardize the image distribution; σ is the standard deviation of the pixel values; I'(x, y) is the pixel value after normalization.

[0023] A further technical solution of the present invention is that the step S20 includes:

[0024] Step S201: Construct a semantic guidance network and combine it with a diffusion denoising network to generate a reconstructed image.

[0025] A further technical solution of the present invention is that the step S201 includes: Use a pre-trained pixel space autoencoder to perform latent space characterization on the input image. Among them, the encoder mapping formula is: z = f enc (I);

[0026] Where is the feature vector in the latent space, representing the high-dimensional characterization of the lens image; f enc is the encoder, used to map the original pixels to the latent space; I is the input lens image.

[0027] A further technical solution of the present invention is that in the step S201, a semantic guidance network g(z) and a denoising diffusion model h(z) are introduced, and the joint optimization objective function is:

[0028] Among them, L rec is the reconstruction loss, which measures the similarity between the reconstructed image and the original input image; L diff is the regularization term of the denoising diffusion model, which improves the stability of image generation; λ is the weight parameter of the regularization term in the loss function, which is used to balance the contributions of the two losses.

[0029] A further technical solution of the present invention is that in the step S201, the robustness of defect reconstruction at different scales is improved through a multi-scale feature fusion module, where the multi-scale feature fusion module is expressed as:

[0030]

[0031] Among them, F fused is the fused multi-scale feature; F i is the feature representation extracted at the i-th scale; α i is the weight coefficient corresponding to the i-th feature, which is dynamically adjusted through training; N is the total number of scales used in feature fusion;

[0032] In the step S201, the skip connection mechanism is combined to fuse the high-scale semantic features and the low-scale detail features to achieve fine reconstruction of the defect area.

[0033] A further technical solution of the present invention is that in the step S201, the steps of generating the reconstructed image include:

[0034] Mapping the features in the latent space back to the pixel space:

[0035] I rec = f dec (F fused );

[0036] Among them, f dec represents the decoder, I rec is the reconstructed lens image; I rec is the reconstructed image; f dec is the decoder, which restores the fused latent space features to the pixel space.

[0037] A further technical solution of the present invention is that in the step S30, the steps of feature extraction and similarity calculation include:

[0038] Step S301, using a pre-trained network to extract the multi-scale features of the input image I and the reconstructed image I rec :

[0039] Among them, is the feature extracted from the input lens image at the l-th layer; The features extracted at the l-th layer for reconstructing the lens image; The l-th layer feature extractor of the ResNet network;

[0040] Step S302, calculate the cosine similarity between features: Anomaly score: A(x,y) = 1 - S (l) (x,y);

[0041] Among them, <·,·> is the inner product operation, used to measure the similarity of feature vectors; ‖·‖ is the Euclidean norm of the vector; S (l) is the feature similarity between the input image and the reconstructed image at the l-th layer; A(x,y) is the anomaly score of the pixel point (x,y), and the higher the value, the greater the degree of anomaly.

[0042] A further technical solution of the present invention is that the step of generating an abnormal defect distribution map and performing defect localization in step S30 includes:

[0043] Integrate the multi-scale anomaly scores into the anomaly distribution map A(x,y), and realize the localization of the defect area through the binarization method:

[0044]

[0045] Among them, τ is the threshold, used to distinguish normal and abnormal points; D(x,y) is the classification result of the pixel point, 1 represents the abnormal area, and 0 represents the normal area.

[0046] To achieve the above object, the present invention also proposes a lens defect detection system based on region segmentation and semantic guidance. The system includes a memory, a processor, and a lens defect detection program based on region segmentation and semantic guidance stored on the processor. When the lens defect detection program based on region segmentation and semantic guidance is run by the processor, it executes the steps of the method described above.

[0047] The beneficial effects of the lens defect detection method and system based on region segmentation and semantic guidance of the present invention are:

[0048] Through the above technical solutions, the present invention preprocesses the lens defect image based on the region segmentation preprocessing method; performs feature fusion on the preprocessed image based on the semantic network and spatial perception feature fusion to reconstruct the lens image, maps the features in the latent space back to the pixel space; extracts features and calculates similarity for the reconstructed lens image to generate an abnormal defect distribution map and perform defect localization; uses a classification model to classify the detected defect regions and output lens defect information, where the defect information includes the defect location, type, and severity. First, preprocessing the lens defect image solves the problem of regular differences in the background and texture features of different regions due to different materials. The present invention makes full use of the geometric information and semantic features of the lens to accurately detect the defect regions, and uses the semantic guidance network and multi-scale feature fusion to improve the reconstruction ability of the defect regions. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on the structures shown in these drawings.

[0050] Figure 1 is a schematic flowchart of a preferred embodiment of the lens defect detection method based on region segmentation and semantic guidance of the present invention;

[0051] Figure 2 is a schematic overall flowchart of the lens defect detection method based on region segmentation and semantic guidance of the present invention;

[0052] Figure 3 is a schematic diagram of region segmentation preprocessing;

[0053] Figure 4 is a cutting diagram of ROI_1 area;

[0054] Figure 5 is a cutting diagram of ROI_2 area;

[0055] Figure 6 is a cutting diagram of ROI_3 area.

[0056] The realization, functional features, and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0058] The present invention proposes a lens defect detection method based on region segmentation and semantic guidance. The present invention is a lens multi-ring partition detection scheme based on the Hough circle transform for a lens, which is beneficial to improving the efficiency and accuracy of lens defect detection.

[0059] Please refer to Figure 1 to Figure [X], and the preferred embodiment of the lens defect detection method based on region segmentation and semantic guidance of the present invention includes the following steps:

[0060] Step S10, preprocess the lens defect image based on the region segmentation preprocessing method.

[0061] In this embodiment, by preprocessing the lens defect image, the problem of regular differences in background and texture features in different regions due to different materials is solved.

[0062] Step S20, perform feature fusion on the preprocessed image based on the fusion of semantic network and spatial perception features, reconstruct the lens image, and map the features in the latent space back to the pixel space.

[0063] Step S30, extract features and calculate similarity for the reconstructed lens image, generate an abnormal defect distribution map, and perform defect localization.

[0064] Step S40, use a classification model to classify the detected defect regions, and output lens defect information, where the defect information includes the defect location, type, and severity.

[0065] This embodiment makes full use of the geometric information and semantic features of the lens to accurately detect the defect regions, and uses a semantic guidance network and multi-scale feature fusion to improve the reconstruction ability of the defect regions.

[0066] In this embodiment, the specific steps of step S10 include:

[0067] Step S101, collect high-resolution defect images of the lens through an industrial camera to ensure that the detailed textures of different regions are clearly presented;

[0068] Step S102, use the region segmentation method to perform image cutting and rectangular filling of the lens.

[0069] The specific steps of step S102 include:

[0070] Use the Hough circle transform to detect the center coordinates of the smallest circle in the image;

[0071] Cut the image according to the center coordinates and the inner and outer diameters of multiple rings set, and perform partitioning. Since the production processes and lens sizes of products with the same mold number are the same, the inner / outer diameter parameters for dividing the rings are flexibly and dynamically set according to the template image;

[0072] Cut the ring or circle and return the minimum bounding rectangle;

[0073] Step S103, perform normalization and scaling processing on the lens image, and adjust the image size to unify the model input format. Among them, the normalization formula is:

[0074]

[0075] Among them, I(x,y) is the pixel value of the input image; μ is the mean of all pixels, used to standardize the image distribution; σ is the standard deviation of the pixel values; I'(x,y) is the normalized pixel value.

[0076] Furthermore, in this embodiment, the step S20 includes:

[0077] Step S201, construct a semantic guidance network and combine it with a diffusion denoising network to generate a reconstructed image.

[0078] The step S201 specifically includes: using a pre-trained pixel space auto-encoder (such as VQVAE, MAE) to perform latent space representation on the input image. Among them, the encoder mapping formula is: z = f enc (I);

[0079] Among them, is the feature vector in the latent space, representing the high-dimensional representation of the lens image; f enc is the encoder, used to map the original pixels to the latent space; I is the input lens image.

[0080] In the step S201, a semantic guidance network g(z) and a denoising diffusion model h(z) are introduced, and the joint optimization objective function is:

[0081] Among them, L rec is the reconstruction loss, measuring the similarity between the reconstructed image and the original input image; L diff is the regularization term of the denoising diffusion model, improving the stability of image generation; λ is the weight parameter of the regularization term in the loss function, used to balance the contributions of the two losses.

[0082] Furthermore, in this embodiment, in step S201, a multi-scale feature fusion module is used to enhance the robustness of defect reconstruction at different scales. The multi-scale feature fusion module is expressed as:

[0083]

[0084] where F fused is the fused multi-scale feature; F i is the feature representation extracted at the i-th scale; α i is the weight coefficient corresponding to the i-th feature, which is dynamically adjusted through training; N is the total number of scales used in feature fusion;

[0085] In this embodiment, in step S201, in combination with the skip connection mechanism, high-scale semantic features and low-scale detail features are fused to achieve fine reconstruction of the defect area.

[0086] In this embodiment, the steps of generating the reconstructed image in step S201 include:

[0087] Mapping the features in the latent space back to the pixel space:

[0088] I rec = f dec (F fused );

[0089] where f dec represents the decoder, I rec is the reconstructed lens image; I rec is the reconstructed image; f dec is the decoder, which restores the fused latent space features to the pixel space.

[0090] In this embodiment, the steps of feature extraction and similarity calculation in step S30 include:

[0091] Step S301, using a pre-trained network (such as ResNet50) to extract the multi-scale features of the input image I and the reconstructed image I rec :

[0092] where is the feature extracted from the input lens image at the l-th layer; is the feature extracted from the reconstructed lens image at the l-th layer; is the l-th layer feature extractor of the ResNet network;

[0093] Step S302, calculating the cosine similarity between the features: Anomaly score: A(x, y) = 1 - S (l) (x, y);

[0094] wherein, <·,·> is an inner product operation used to measure the similarity of feature vectors; ‖·‖ is the Euclidean norm of a vector; S (l) is the feature similarity between the input image and the reconstructed image at the l-th layer; A(x, y) is the anomaly score of the pixel point (x, y), and the higher the value, the greater the degree of anomaly.

[0095] Further, in this embodiment, the step S30 of generating an abnormal defect distribution map and performing defect localization includes:

[0096] Integrate the multi-scale anomaly scores into the anomaly distribution map A(x, y), and realize the localization of the defect area through a binarization method:

[0097]

[0098] where τ is a threshold for distinguishing normal and abnormal points; D(x, y) is the classification result of the pixel point, 1 represents the abnormal area, and 0 represents the normal area.

[0099] The beneficial effects of the lens defect detection method based on region segmentation and semantic guidance of the present invention are:

[0100] Through the above technical solutions, the present invention preprocesses the lens defect image based on the region segmentation preprocessing method; performs feature fusion on the preprocessed image based on the semantic network and spatial perception feature fusion, reconstructs the lens image, and maps the features in the latent space back to the pixel space; performs feature extraction and similarity calculation on the reconstructed lens image, generates an abnormal defect distribution map, and performs defect localization; uses a classification model to classify the detected defect area and output lens defect information, where the defect information includes the defect location, type, and severity. First, preprocessing the lens defect image solves the problem of regular differences in background and texture features in different regions due to different materials. The present invention makes full use of the geometric information and semantic features of the lens to accurately detect the defect area, and uses the semantic guidance network and multi-scale feature fusion to improve the reconstruction ability of the defect area.

[0101] To achieve the above object, the present invention also proposes a lens defect detection system based on region segmentation and semantic guidance. The system includes a memory, a processor, and a lens defect detection program based on region segmentation and semantic guidance stored on the processor. When the lens defect detection program based on region segmentation and semantic guidance is run by the processor, it executes the steps of the method described above, which will not be elaborated here.

[0102] The above are only the preferred embodiments of the present invention, and do not thereby limit the patent scope of the present invention. Any equivalent structural transformation made under the concept of the present invention by using the content of the specification and drawings of the present invention, or any direct / indirect application in other related technical fields is included within the patent protection scope of the present invention.

Claims

1. A lens defect detection method based on region segmentation and semantic guidance, characterized in that The method includes the following steps: Step S10, preprocessing the lens defect image based on a region segmentation preprocessing method; Step S20, performing feature fusion on the preprocessed image based on semantic network and spatial perception feature fusion, reconstructing the lens image, and mapping the features in the latent space back to the pixel space; Step S30, performing feature extraction and similarity calculation on the reconstructed lens image, generating an abnormal defect distribution map, and performing defect localization; Step S40, using a classification model to classify the detected defect regions, and outputting lens defect information, where the defect information includes the defect position, type, and severity.

2. The method for detecting lens defects based on region segmentation and semantic guidance according to claim 1, characterized in that, The step S10 includes: Step S101, collecting a high-resolution defect image of the lens through an industrial camera to ensure clear presentation of the detailed textures in different regions; Step S102, using a region segmentation method to perform image cutting and rectangular complementing of the lens: Detecting the center coordinates of the smallest circle in the image using the Hough circle transform; Cutting the image according to the center coordinates and setting the inner and outer diameters of multiple rings for partitioning; Cutting the rings or circles and returning the minimum bounding rectangle; Step S103, performing normalization and scaling processing on the lens image to adjust the image size, where the normalization formula is: where I(x,y) is the pixel value of the input image; μ is the mean of all pixels, used to standardize the image distribution; σ is the standard deviation of the pixel values; I'(x,y) is the normalized pixel value.

3. The method for detecting lens defects based on region segmentation and semantic guidance according to claim 2, wherein The step S20 includes: Step S201, constructing a semantic guidance network and combining it with a diffusion denoising network to generate a reconstructed image.

4. The method for detecting lens defects based on region segmentation and semantic guidance according to claim 3, wherein The step S201 includes: performing latent space representation on the input image using a pre-trained pixel space autoencoder, where the encoder mapping formula is: z = f enc (I); Among them, is the feature vector of the latent space, representing the high-dimensional representation of the lens image; f enc is the encoder for mapping the original pixels to the latent space; I is the input lens image.

5. The method for detecting lens defects based on region segmentation and semantic guidance according to claim 4, wherein In the step S201, a semantic guidance network g(z) and a denoising diffusion model h(z) are introduced, and the objective function is jointly optimized: Among them, L rec is the reconstruction loss, which measures the similarity between the reconstructed image and the original input image; L diff is the regularization term of the denoising diffusion model, which improves the stability of image generation; λ is the weight parameter of the regularization term in the loss function, which is used to balance the contributions of the two losses.

6. The method for detecting lens defects based on region segmentation and semantic guidance according to claim 5, wherein, In the step S201, the robustness of defect reconstruction at different scales is improved through a multi-scale feature fusion module, where the multi-scale feature fusion module is expressed as: Among them, F fused is the fused multi-scale feature; F i is the feature representation extracted at the i-th scale; α i is the weight coefficient corresponding to the i-th feature, which is dynamically adjusted through training; N is the total number of scales used in feature fusion; In the step S201, a skip connection mechanism is combined to fuse high-scale semantic features with low-scale detail features to achieve fine reconstruction of the defect region.

7. The method for detecting lens defects based on region segmentation and semantic guidance according to claim 6, wherein In the step S201, the steps for generating the reconstructed image include: Mapping the features in the latent space back to the pixel space: I rec = f dec (F fused ); Among them, f dec represents a decoder, and I rec is the reconstructed lens image; I rec is the reconstructed image; f dec is the decoder that restores the fused latent space features to the pixel space.

8. The method for detecting lens defects based on region segmentation and semantic guidance according to claim 7, characterized in that, The steps for feature extraction and similarity calculation in the step S30 include: Step S301, use a pre-trained network to extract multi-scale features of the input image I and the reconstructed image I rec : Among them, is the feature extracted from the input lens image at the l-th layer; is the feature extracted from the reconstructed lens image at the l-th layer; is the l-th layer feature extractor of the ResNet network; Step S302, calculate the cosine similarity between features: Anomaly score: A(x, y) = 1 - S (l) (x, y); Among them, <·,·> is the inner product operation, which is used to measure the similarity of feature vectors; ‖·‖ is the Euclidean norm of the vector; S (l) is the feature similarity between the input image and the reconstructed image at the l-th layer; A(x, y) is the anomaly score of the pixel point (x, y), and the higher the value, the greater the degree of anomaly.

9. The method for detecting lens defects based on region segmentation and semantic guidance according to claim 8, wherein, The steps for generating the abnormal defect distribution map and performing defect localization in the step S30 include: Integrating the multi-scale abnormal scores into an abnormal distribution map A(x,y), and achieving defect region localization through a binarization method: where τ is a threshold used to distinguish normal and abnormal points; D(x,y) is the classification result of the pixel point, 1 represents the abnormal region, and 0 represents the normal region.

10. A lens defect detection system based on region segmentation and semantic guidance, the system includes a memory, a processor, and a lens defect detection program based on region segmentation and semantic guidance stored on the processor. When the lens defect detection program based on region segmentation and semantic guidance is run by the processor, it executes the steps of the method according to any one of claims 1 to 9.