Glass lens surface scratch detection method and system

The multi-angle polarization light source array and the near-infrared compensation light source generate an orthogonal polarization composite light field, combined with the asymmetric Gabor filter group and the improved OTSU threshold segmentation method, the error detection and missed detection of glass lens scratch detection under complex coating processes are solved, and high-precision and efficient scratch recognition are achieved.

CN120446165AInactive Publication Date: 2025-08-08NANYANG CITY JINGLIANG OPTICAL TECH CO LTD

Patent Information

Application Number
CN202510515527.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, in complex coating processes or high reflectivity surface scenes, the light transmittance of glass lenses, random reflection noise of coating layer and background texture form multi-scale coupling interference, resulting in high false alarm rate and high leakage detection rate of scratch recognition algorithms, and it is difficult to adapt to inclination, curves or intermittent scratches, affecting detection accuracy and real-time performance.

Method used

A multi-angle polarized light source array and a near-infrared compensation light source are used to generate an orthogonal polarized composite light field, combined with an asymmetric Gabor filter group and an improved OTSU threshold segmentation method, through the Gaussian pyramid multi-scale feature fusion and direction-constrained convolution kernel group, the light source wavelength and convolution kernel offset are dynamically adjusted, and the direction-constrained decoupling module and cascade network model are constructed to achieve accurate detection of scratches.

Benefits of technology

It significantly reduces the error detection rate caused by light field interference, enhances the algorithm's generalization ability of defects in different forms, improves the spatial resolution and real-time nature of scratch detection, and meets the large-scale application needs of high-speed production lines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120446165A_ABST
    Figure CN120446165A_ABST
Patent Text Reader

Abstract

The invention discloses a glass lens surface scratch detection method and system, relates to the technical field of precision optical detection, and aims to solve the problems of scratch false detection, leak detection and poor algorithm adaptability caused by interference fringes, noise coupling and poor form adaptability in a high-reflection / complex coating process scene in the prior art. According to the scheme, an orthogonal polarization state composite light field is generated based on a multi-angle polarization light source array and a near-infrared compensation light source, and candidate regions are extracted through dynamic threshold segmentation and a direction gradient tensor matrix; gaussian pyramid multi-scale feature fusion and refraction angle consistency verification are utilized to eliminate artifact interference; constructing a direction constraint convolution kernel group to decompose scratches and background textures, and dynamically allocating computing resources in combination with a cascade network; feeding back closed-loop calibration light source wavelength and convolution kernel parameters in real time through coating parameters; according to the method, the precision and robustness of high-reflectivity surface scratch detection are remarkably improved, and meanwhile, the requirements for high-resolution image processing and real-time performance in a high-speed production line are balanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of precision optical detection technology, and more particularly to a method and system for detecting scratches on the surface of a glass lens. Background Art

[0002] With the rapid development of precision optical manufacturing, consumer electronics, and the automotive industry, the high-precision surface quality of glass lenses has become a core indicator affecting imaging performance and product reliability. Traditional manual visual inspection is inefficient and highly subjective, making it difficult to meet the large-scale production needs of micron-level scratches. In recent years, the integration of machine vision and artificial intelligence technologies has promoted the progress of automated surface defect detection. However, the light transmittance of glass materials, the complex reflective characteristics of coating processes, and the variability of scratch morphology (such as invisible scratches and intermittent scratches) still pose severe challenges to the accuracy, anti-interference ability, and real-time performance of the detection system.

[0003] Current glass lens surface scratch detection technology mainly relies on a combination of optical imaging and algorithmic analysis. For example, patent CN119379620A is based on a method of color grid and scratch shortest distance screening. It generates a scratch color grid by graying the image, and uses the shortest neighbor distance of color to screen the real scratch area. This method relies on the effectiveness of the image processing algorithm, but may be affected under complex lighting conditions. Another type of method uses laser irradiation and light propagation image analysis, such as described in patent CN119470486A. Through the interaction between laser and glass, light propagation images are collected, and abnormal pixels are analyzed after preprocessing to detect defects. In addition, there are methods based on artificial intelligence, such as described in patent CN119804505A, which use intelligent analysis terminals and shooting equipment to obtain glass images, determine the defect location through feature analysis processing, and evaluate the quality based on the distance between the defect and the center of the glass.

[0004] However, in the above existing technologies, in complex coating processes or high-reflectivity surface scenarios, the transmittance of glass lenses, the random reflection noise of the coating layer, and the background texture form multi-scale coupling interference, making it difficult for traditional grayscale analysis or color grid algorithms to effectively decouple the real scratch features. For example, in areas where the mirror reflectivity exceeds 85% (such as the edge of a high-refractive-index lens), traditional brightfield illumination will produce interference fringes, causing the false alarm rate of the scratch recognition algorithm to soar. In addition, in metal coatings or frosted surfaces, the grayscale distribution of noise and scratches highly overlaps, causing false detection or missed detection. In addition, the color grid technology has limited adaptability to the direction and shape of scratches. It relies on the scratch screening rule of a fixed neighborhood distance and cannot accurately cover the target area in inclined, curved or intermittent scratch scenarios. It needs to rely on manual parameter adjustment to adapt to defects of different shapes, which significantly reduces the robustness of the algorithm. At the same time, existing imaging systems based on lasers or static light sources have difficulty suppressing the interference of complex backgrounds such as impurities inside the glass and edge processing texture on surface scratch detection, resulting in a decrease in defect positioning accuracy. In dynamic detection scenarios, the computing power requirements of high-resolution image processing and multi-feature traversal calculations further exacerbate the contradiction between the real-time performance of the algorithm and detection accuracy, restricting its large-scale application in high-speed production lines. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention discloses a method and system for detecting scratches on the surface of a glass lens, aiming to solve the problems in the background technology.

[0006] In order to achieve the above technical effects, the present invention adopts the following technical solutions:

[0007] A method for detecting scratches on a glass lens surface, comprising:

[0008] Step 1: Based on the multi-angle polarized light source array, a four-directional polarizer and a near-infrared compensation light source are integrated to generate a composite light field with orthogonal polarization states. The compensation light spot is synchronously projected onto the coating layer through a time-sharing trigger mechanism, and a four-channel polarized reflection image sequence and a coating interference compensation map are output;

[0009] Step 2: Based on the four-channel polarized reflectance image sequence, a directional gradient tensor matrix is generated through an asymmetric Gabor filter bank. Combined with the noise distribution of the coating interference compensation image, the threshold segmentation parameters are dynamically adjusted through the improved OTSU threshold segmentation method to output the binary scratch candidate area and directional coding matrix;

[0010] Step 3: Input the binary scratch candidate area into the Gaussian pyramid decomposition module, extract multi-scale morphological features layer by layer, fuse low-frequency continuous features with high-frequency detail features through a cross-layer attention weighting mechanism, use a refraction angle consistency verification algorithm to eliminate artifact interference, and output a fine-screened scratch area map;

[0011] Step 4: Based on the directional coding matrix and the finely screened scratch area map, a directional constrained convolution kernel group is constructed, dynamic convolution operations are performed within the restricted directional interval, and scratches and background textures are quantized and decomposed through a multi-scale texture decomposition module, and a decoupled feature map and spatial orthogonalized code are output;

[0012] Step 5: Input the decoupled feature map into the cascade network model, evaluate the confidence based on the directional distribution heat map, dynamically allocate FP16 quantization computing resources to high-probability areas, and output the scratch topology structure containing geometric parameters and 3D depth information;

[0013] Step 6: Obtain coating process parameters, including film thickness and refractive index, in real time through the CAT bus, dynamically adjust the compensation light source wavelength of step 1, perform reverse projection error analysis on the output of step 5 through a standard sample, close-loop calibrate the light source incident angle of step 1 and the convolution kernel offset of step 4, and update the light synergy parameter library.

[0014] As a further technical solution of the present invention, a glass lens surface scratch detection system includes: a multimodal optical acquisition module, a gradient domain dynamic processing module, a multi-scale feature fusion module, a directional constraint decoupling module, a three-dimensional topology modeling module and a dynamic closed-loop optimization module;

[0015] The multimodal optical acquisition module is used to integrate a four-directional polarized light source array and a near-infrared compensation light source, generate a composite light field with orthogonal polarization states through time-sharing triggering, synchronously project the compensation light spot onto the coating surface, and output a four-channel polarized reflection image sequence and a coating interference compensation map;

[0016] The gradient domain dynamic processing module is used to perform multi-directional gradient convolution on the four-channel polarized reflectance image sequence based on the asymmetric Gabor filter bank, dynamically optimize the segmentation threshold of the improved OTSU algorithm in combination with the noise distribution characteristics of the coating interference compensation image, and output the binary scratch candidate area and the scratch direction encoding matrix;

[0017] The multi-scale feature fusion module is used to input the binary scratch candidate area into the Gaussian pyramid decomposition network, extract multi-scale morphological features in layers, fuse low-frequency continuous features with high-frequency detail features through a cross-layer attention weighting mechanism, and eliminate artifact interference based on the refraction angle consistency verification algorithm, and output a high-confidence scratch area map after fine screening;

[0018] The directional constraint decoupling module is used to construct a directional constraint dynamic convolution kernel group based on the directional coding matrix and the finely screened scratch area map, perform multi-scale texture orthogonal decomposition within the restricted directional range, quantize and separate scratch features and base texture noise, and output a decoupled feature map and a spatial orthogonalized code;

[0019] The three-dimensional topology modeling module is used to input the decoupled feature map into the cascade network model, dynamically allocate FP16 quantization computing resources based on the directional distribution heat map, and fuse the spatial orthogonalization coding to generate a scratch topology structure containing three-dimensional depth information and geometric parameters;

[0020] The dynamic closed-loop optimization module is used to obtain the coating thickness and refractive index parameters in real time through the CAT bus, perform back-projection error analysis on the scratch topology structure in combination with a standard template, dynamically adjust the compensation light source wavelength of the multimodal optical acquisition module and the convolution kernel offset of the directional constraint decoupling module, and update the light field collaborative parameter library.

[0021] Based on the above technical solutions, the positive and beneficial effects of the present invention are:

[0022] 1. By using a multi-angle polarized light source array and orthogonal polarization state composite light field technology, the wavelength and incident angle of the light source are dynamically adjusted to directly weaken the interference fringe generation conditions in the mirror reflection area. Combined with real-time feedback from the coating interference compensation map, the light wave interference effect on high-reflectivity (>85%) surfaces is suppressed from the source of optical imaging, significantly reducing the false detection rate caused by light field interference and ensuring the detection reliability of complex areas such as the edges of high-refractive-index lenses. The dynamic noise suppression mechanism of the directional gradient tensor calculation of the asymmetric Gabor filter group and the improved OTSU algorithm is used. Through local variance-driven adaptive window segmentation, the grayscale overlapping areas of noise and scratches on metal coatings or frosted surfaces are accurately separated. Combined with the noise mask weighted decision, background texture and real defects are effectively distinguished, reducing misjudgments and missed detections caused by grayscale confusion.

[0023] 2. Multi-scale texture decomposition technology based on a directional-constrained convolution kernel group breaks through the fixed neighborhood limitations of traditional color-rendering grids. By dynamically adapting the scratch's trajectory to a restricted directional interval, it covers the complete morphological characteristics of inclined, curved, or intermittent scratches. Combined with a cross-layer attention weighting mechanism, it adaptively fuses multi-scale features, eliminating reliance on manual parameter adjustment and enhancing the algorithm's generalization capabilities for defects of varying morphologies. Through multi-scale texture orthogonal decomposition and a refraction angle consistency verification algorithm, it quantitatively separates complex background noise such as internal impurities and edge processing textures in glass. Combining 3D point cloud generation technology with depth information and geometric parameters, it accurately locates the spatial distribution of surface scratches, avoiding positioning deviations caused by background interference or internal defects and improving the spatial resolution of defect recognition.

[0024] 3. Based on the dynamic allocation mechanism of FP16 quantized computing resources, computing resources are allocated preferentially to high-confidence areas, optimizing the computational efficiency of high-resolution image processing. Combining multi-physics field coupling noise suppression and compressed sensing reconstruction technology, while ensuring sub-micron detection accuracy, the computing power requirements are significantly reduced, meeting the real-time requirements of high-speed production lines and promoting the implementation of large-scale industrial applications. And through the real-time feedback of coating process parameters and the dynamic calibration mechanism driven by back-projection error, a collaborative optimization closed loop of the detection system and the manufacturing process is formed. The light source parameters and convolution kernel offset are adaptively adjusted to compensate in real time for the impact of process disturbances such as film thickness fluctuations and refractive index changes on the detection results, thereby improving the stability and adaptability of the system in complex coating scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive work, among which:

[0026] Figure 1 This is a structural diagram of a method for detecting scratches on a glass lens surface according to the present invention;

[0027] Figure 2 This is a working principle diagram of the glass lens surface scratch detection system of the present invention;

[0028] Figure 3 This is a working principle diagram of step 1 of the present invention;

[0029] Figure 4 Schematic diagram of the working principle of the asymmetric Gabor filter bank of the present invention;

[0030] Figure 5 This is a structural framework diagram of the multi-scale texture decomposition module of the present invention;

[0031] Figure 6 It is a structural framework diagram of the cascade network model of the present invention. DETAILED DESCRIPTION

[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0033] In an embodiment, a method for detecting scratches on a glass lens surface is provided. Figure 1 Shown, including:

[0034] Step 1: Based on the multi-angle polarized light source array, a four-directional polarizer and a near-infrared compensation light source are integrated to generate a composite light field with orthogonal polarization states. The compensation light spot is synchronously projected onto the coating layer through a time-sharing trigger mechanism, and a four-channel polarized reflection image sequence and a coating interference compensation map are output;

[0035] Step 2: Based on the four-channel polarized reflectance image sequence, a directional gradient tensor matrix is generated through an asymmetric Gabor filter bank. Combined with the noise distribution of the coating interference compensation image, the threshold segmentation parameters are dynamically adjusted through the improved OTSU threshold segmentation method to output the binary scratch candidate area and directional coding matrix;

[0036] Step 3: Input the binary scratch candidate area into the Gaussian pyramid decomposition module, extract multi-scale morphological features layer by layer, fuse low-frequency continuous features with high-frequency detail features through a cross-layer attention weighting mechanism, use a refraction angle consistency verification algorithm to eliminate artifact interference, and output a fine-screened scratch area map;

[0037] Step 4: Based on the directional coding matrix and the finely screened scratch area map, a directional constrained convolution kernel group is constructed, dynamic convolution operations are performed within the restricted directional interval, and scratches and background textures are quantized and decomposed through a multi-scale texture decomposition module, and a decoupled feature map and spatial orthogonalized code are output;

[0038] Step 5: Input the decoupled feature map into the cascade network model, evaluate the confidence based on the directional distribution heat map, dynamically allocate FP16 quantization computing resources to high-probability areas, and output the scratch topology structure containing geometric parameters and 3D depth information;

[0039] Step 6: Obtain coating process parameters, including film thickness and refractive index, in real time through the CAT bus, dynamically adjust the compensation light source wavelength of step 1, perform reverse projection error analysis on the output of step 5 through a standard sample, close-loop calibrate the light source incident angle of step 1 and the convolution kernel offset of step 4, and update the light synergy parameter library.

[0040] In step 1 of the above embodiment, Figure 3As shown, the working principle of step 1 is as follows: based on the wavelength polarization direction coding protocol, polarization angle and wavelength parameters are assigned to each LED unit in the multi-angle polarized light source array, wherein the visible light polarized light source group uses a four-way polarizer switching mechanism to output linearly polarized light with polarization angles of 0°, 45°, 90°, and 135° and wavelengths of 450nm and 650nm respectively; and synchronously triggers the 850nm near-infrared compensation light source, and projects the compensation light spot to the interface between the coating layer and the substrate through the polarization beam splitter. The industrial camera captures a four-channel polarized reflection image sequence based on the synchronous trigger signal, and each channel is The reflected light intensity distribution of a single polarization angle and wavelength combination is obtained; then, the polarization component parallel to the incident plane and the polarization component perpendicular to the incident plane are separated by a four-channel polarization solution algorithm, and the polarized reflection image sequence of four time-sharing exposures is dynamically fused; in response to the interference of coating layer thickness fluctuations, the coating process parameters are obtained in real time through the CAT bus, and a closed-loop feedback control algorithm is used to drive the rotating motor to adjust the azimuth angle of the polarizer to the extinction position to reduce the mirror reflection light intensity; finally, the grayscale gradient covariance analysis method is used to calculate the pixel intensity variance matrix of the near-infrared compensation area and the non-compensation area to generate the coating interference compensation map.

[0041] The multi-angle polarized light source array utilizes a wavelength-polarization encoding protocol, arranging visible light (450nm and 650nm) and near-infrared light (850nm) light source units at specific spatial angles. The polarization angle of each LED unit is precisely controlled by a four-way polarizer switching mechanism. The multi-wavelength design in the visible light band enhances the differences in the scattering response of different material coatings to scratches, while the penetrability of near-infrared light allows it to bypass the surface coating and reach the substrate interface, separating the surface scratch and internal impurity reflection signals through the optical path difference between the substrate and the coating. The four-way polarizer switching mechanism switches between four polarization angles of 0°, 45°, 90°, and 135° in a time-sharing manner, covering the complete Stokes parameter space of polarization states. This allows scratch edge gradients in any direction to be captured by at least two orthogonal polarization channels, avoiding feature loss caused by a single polarization state. The time-sharing trigger mechanism synchronizes the light source pulses with the exposure timing of the industrial camera through FPGA hardware, ensuring that the reflected light intensity of each polarization angle / wavelength combination is independently imaged within the time window, eliminating multi-channel crosstalk. The polarization beam splitter decomposes the reflected light into polarization components parallel (p) and perpendicular (s) to the incident plane. Combined with the polarization solution algorithm of the four-channel image sequence, the Jones matrix model is used to reconstruct the polarization reflectivity distribution of each point on the coating surface, thereby suppressing the artifact signal dominated by specular reflection.

[0042] Coating interference compensation acquires coating thickness and refractive index parameters in real time via the CAT bus. Based on thin-film interference theory, the incident angle adjustment of the near-infrared compensation light source is dynamically calculated, ensuring that the reflection path of the compensation spot forms destructive interference with the light reflected from the coating-substrate interface, thereby reducing random noise caused by coating fluctuations. Closed-loop feedback control utilizes a PID algorithm to drive a rotary motor to fine-tune the polarizer azimuth to the extinction position (the point of minimum specular reflection intensity). Combined with grayscale gradient covariance analysis, the pixel intensity variance matrices of compensated and non-compensated areas are compared to quantify the spatial distribution of coating interference noise and generate a compensation map for subsequent noise modeling and dynamic suppression.

[0043] In one implementation, the multi-angle polarized light source array consisted of 24 LED units arranged in a ring, each containing 450nm blue and 650nm red LEDs. A four-way polarizer switching module (driven by piezoelectric ceramics with a rotation accuracy of ±0.1°) was integrated on the outside. The near-infrared compensation light source used an 850nm LED array, which was directed onto the lens surface via a polarizing beam splitter (extinction ratio >1000:1). The industrial camera used a global shutter CMOS sensor (120fps) and received a synchronous trigger signal from the light source controller via a GPIO interface, with a trigger delay of less than 1μs. A rotary motor (0.01° step angle) was connected to an industrial computer via an RS485 interface, receiving coating process parameters in real time and adjusting the polarizer angle. At the software level, the time-sharing trigger sequence was set to sequentially activate the four visible light source groups at polarization angles of 0°, 45°, 90°, and 135°, with a single exposure time of 2ms. The near-infrared light source was synchronously triggered at the end of each cycle for a 1ms exposure. The four-channel polarization solution algorithm converts the four-frame image into p components (I p ) and s component (I s ) intensity ratio, the formula is I p / I s =(I0-I 90 ) / (I 45 -I 135 The closed-loop feedback control algorithm calculates the optimal extinction angle θ = arcsin (nλ / (4h)) based on the coating thickness h and the refractive index n, and drives the motor to rotate within the range of θ ± 0.5° to iteratively search for the optimal value. In the grayscale gradient covariance analysis, the compensation map generation module traverses the near-infrared image with a 15×15 pixel window to calculate the local variance σ 2 =∑(I(x,y)-μ) 2 / N, and pass the threshold σ 2 A high interference noise area is determined when the value is >120, and the pixel intensity in this area is weighted attenuated in the visible light channel (attenuation coefficient 0.3-0.7).

[0044] During implementation, this step significantly suppresses mirror reflection and coating interference noise on highly reflective surfaces through polarization state encoding and multi-wavelength collaborative illumination, thereby enhancing the ability to extract scattered signals in scratched areas; the dynamic extinction angle adjustment mechanism effectively adapts to changes in different coating process parameters, solving the adaptability limitations of traditional fixed-angle polarization detection; near-infrared compensation and grayscale covariance analysis achieve accurate modeling and spatial domain suppression of background noise, improve the signal-to-noise ratio of scratch features, and provide high-quality polarization image input for subsequent processing. The overall technical solution has higher detection accuracy and robustness than existing methods in highly reflective and complex coating scenarios.

[0045] In step 2 of the above embodiment, the asymmetric Gabor filter bank is configured with differentiated aspect ratios and phase offset parameters along 12 directions, wherein the ratio of the major axis to the minor axis of the horizontal and vertical kernels is 3:1, the ratio of the diagonal kernels is 2.5:1, and the ratio of the major axis to the minor axis of the intermediate direction is 2.8:1; Figure 4 As shown, the working principle of the asymmetric Gabor filter bank is:

[0046] Through multi-directional convolution kernel, each pixel in the input polarized reflectance image sequence is parallelized to extract the gradient amplitude phase response characteristics of each pixel in the anisotropic frequency domain space, and generate 12-directional gradient feature vectors;

[0047] During the initialization phase of the asymmetric Gabor filter bank, a dynamic bandwidth adjustment mechanism is used to adaptively adjust the half-peak bandwidth parameters of the kernel function in each direction according to the noise distribution characteristics of the coating interference compensation map;

[0048] During the convolution process, the real and imaginary components of the orthogonal Gabor are asymmetrically weighted and fused using the directional weight factor, and the amplitude squared sum phase difference compensation algorithm is used to enhance the signal-to-noise ratio of the scratch edge directional response.

[0049] Based on the 12 directional gradient feature vectors, a directional gradient tensor matrix that fuses directional coding and intensity distribution is constructed through a tensor aggregation method.

[0050] The working method of the improved OTSU threshold segmentation method is:

[0051] S1, coating interference compensation diagram C map The noise distribution is input and the noise variance σ in the neighborhood of each pixel (x, y) is calculated. 2 (x,y), if σ is within the current window 2 (x, y) exceeds the preset noise variance threshold η, the window size adaptive reduction mechanism is triggered, and the window size is adjusted to the original size. Generate dynamic window W dyn ;

[0052] S2, in the dynamic window W dyn In the , the weight factor α is introduced to optimize the inter-class variance to avoid over-segmentation in high noise areas. The optimized inter-class variance is defined as:

[0053]

[0054] In formula (1), α is the weight factor, which is defined as α = 1 / (1 + σ 2 ), used to balance global and local feature contributions; μ T is the global grayscale mean, which is used to preserve the image brightness distribution characteristics; μ0 and μ1 are the grayscale means of the foreground and background in the dynamic window respectively; ω0 and ω1 are the pixel weights of scratches and base respectively;

[0055] S3, based on the coating interference compensation diagram C map Generate noise probability matrix N map (x,y), where N map (x,y)=C map (x,y) / max(C map ), define the noise suppression coefficient β=1-N map (x, y), the segmentation threshold of the high-probability noise area is shifted toward the background class by improving the segmentation decision function. The formula expression of the improved segmentation decision function is:

[0056]

[0057] S4, perform Gaussian pyramid downsampling on the directional gradient tensor matrix, and calculate the initial threshold T at the coarse scale layer of 1 / 4 resolution coarse , and update it layer by layer upwards. The formula is:

[0058]

[0059] In formula (3), λ represents the gradient compensation factor, which is defined as OGM fine Represents the directional gradient tensor matrix at a fine scale, which contains the gradient amplitude and direction information at a fine scale and is used to preserve the edge detail features of the scratch;

[0060] S5. Calculate the directional discreteness θ of the pixels in the candidate area by combining the directional encoding matrix std , the calculation formula is:

[0061] θ std =std(DCM(x,y)) (4)

[0062] In formula (4), DCM(x,y) represents the direction encoding matrix of the pixel point (x,y), which is used to record the maximum gradient response direction of each pixel point; syd represents the standard deviation operation, which is used to quantify the discrete degree of the pixel direction angle in the scratch candidate area; in θ std When it is greater than 15°, it is determined to be artifact interference and removed, and finally the binary scratch area and direction code are output.

[0063] In the asymmetric Gabor filter bank, multi-directional convolution kernels are uniformly spaced (15° steps) across 12 directions, covering a range from 0° to 165°. Differentiated aspect ratios (3:1 horizontal / vertical, 2.5:1 diagonal, and 2.8:1 intermediate) are optimized based on the statistical characteristics of scratch morphology. The major axis extends along the scratch extension to enhance edge response, while the minor axis compresses background noise frequency domain energy. A dynamic bandwidth adjustment mechanism adjusts the frequency domain coverage of each directional kernel through feedback control of the half-maximum bandwidth (FWHM) based on the noise power spectral density distribution of the coating interference compensation pattern. The bandwidth is increased in the dominant noise frequency region to suppress high-frequency interference, while the bandwidth is compressed in the scratch signature frequency region to improve signal focusing. Asymmetric weighted fusion uses directional weighting factors to unbalance the real (symmetric) and imaginary (antisymmetric) Gabor components. The real weighting emphasizes edge strength accumulation, while the imaginary weighting enhances phase jump detection. A combination of amplitude squared and phase difference compensation algorithms (the compensation factor is dynamically calculated from the noise covariance matrix) eliminates phase offset interference caused by coating interference. The tensor aggregation method constructs the 12-directional gradient eigenvectors into a second-order structure tensor matrix, obtains the main direction energy and anisotropy coefficient through eigenvalue decomposition, quantifies the directional consistency and continuity of the scratch edge, and provides a directional encoding prior for subsequent segmentation.

[0064] The core of the improved OTSU threshold segmentation method lies in the fusion and optimization of dynamically adaptive coating interference noise characteristics and multi-scale gradient features. In the dynamic window adjustment mechanism, a preset threshold triggers adaptive window size reduction based on the local statistical characteristics of noise variance. This principle stems from the negative correlation between local variance and image texture complexity: high-frequency components in high-noise regions significantly increase local variance. In this case, reducing the window can suppress the propagation of noise interference while maintaining spatial continuity through subsampling of the neighborhood grayscale mean. The essence of optimizing inter-class variance with a weight factor is to balance the conflict between local segmentation and overall brightness distribution by introducing the global grayscale mean as a regularization term. The construction of the weight factor η follows the nonlinear attenuation of the inter-class variance's sensitivity to noise. Its mathematical expression implicitly incorporates the eigendecomposition results of the covariance matrix between the noise variance and the grayscale distribution. Inter-class separability is optimized through the Lagrange multiplier method. The design of the noise suppression coefficient λ incorporates the probability distribution characteristics of the coating interference compensation map. Using Bayesian decision theory, the threshold offset in high-noise areas is modeled as the negative log-likelihood of the background class conditional probability, effectively reducing missegmentation caused by salt and pepper noise. The Gaussian pyramid multi-scale optimization strategy, based on scale space theory, introduces a gradient compensation factor γ when calculating the initial threshold at the coarse-scale layer. Its physical meaning is the normalized value of the covariance matrix trace of the directional gradient tensor at the fine-scale layer, ensuring that the sub-pixel edge topology is preserved when transferring thresholds across scales. The directional dispersion criterion utilizes the principal component analysis results of the directional encoding matrix to quantify the anisotropic characteristics of the scratch area through the standard deviation. The 15° threshold is set based on the confidence interval of directional consistency under the Rayleigh distribution, effectively filtering out pseudo-edge responses caused by random noise.

[0065] In one implementation, an asymmetric Gabor filter bank was implemented using the parallel architecture of a GPU (NVIDIA A100). Four-channel polarized reflectance images were fed into the GPU memory via the PCIe 4.0 bus. 12-directional convolution operations were performed in parallel for each pixel using a CUDA kernel function, with kernel parameters stored in the FPGA's on-chip RAM. A dynamic bandwidth adjustment module interacted with the coating interference compensation image data via an SPI interface, updating the kernel function's frequency domain parameters in real time. Dynamic window adjustment for the improved OTSU threshold segmentation was implemented using the FPGA. The initial window size was 16×16 pixels, and the noise variance threshold was set to 0.1. Upon triggering, the window was switched to an 8×8 window, with bilinear interpolation for smooth transitions. The inter-class variance optimization module used an integral graph to accelerate the calculation of the local mean on the GPU. The weight factor β was calculated using a sigmoid function (β = 1 / (1 + e^{-10(σ 2The noise suppression coefficient γ is dynamically generated using a precomputed LUT table. Multi-scale thresholding is performed using OpenCV to construct a Gaussian pyramid (standard deviation 1.0). The initial threshold of the coarse-scale layer is transferred to the fine layer via bicubic interpolation, and the gradient compensation factor λ is set to 1.8. Directional dispersion verification uses the SIMD instruction set to parallelize the calculation of the directional angle standard deviation, and areas exceeding the standard are removed using a bit mask. The software process is integrated into the ROS node, and the input is a four-channel image (450nm / 650nm polarization image, 850nm compensation image). The output is a binary scratch area and directional encoding matrix. The processing frame rate is 30fps (4K resolution), the GPU memory usage is ≤6GB, and the FPGA logic resource utilization is ≤75%.

[0066] During implementation, the hardware required for this step includes the following: a four-channel polarization camera (FLIR BFS-PGE-16S2C-CS, supporting 4096×2160 resolution and 12-bit depth), a circular LED polarization light source (wavelength 450-650nm, polarization degree >95%, coaxial ring layout), a high-precision displacement stage (PI H-811, repeatability ±0.1μm, XYZ three-axis linkage), an industrial computer (equipped with an NVIDIA A6000 GPU, 48GB of video memory, and a PCIe 4.0 interface), and an aspherical microscope objective lens (Mitutoyo M Plan Apo 20×, NA = 0.42, working distance 20mm). During implementation, the light source and camera are coaxially mounted, with the polarization direction forming a 45° angle with the lens surface normal. The displacement stage carries the lens sample and is directly connected to the GPU via a PCIe interface for real-time data transmission. A beam splitter prism is placed between the objective lens and the camera to separate the four-channel polarization light signals.

[0067] During the design test, to verify the superiority of step 2 (Group A) over the traditional symmetric Gabor filter (isotropic kernel, fixed aspect ratio of 1.5:1) plus standard OTSU segmentation (Group B) on complex coated surfaces, 10 coated glass lenses (scratch length 50-200μm, width 5-20μm) were selected. Each group underwent five repeated tests: a polarized light source was projected at four angles (0°, 45°, 90°, and 135°), and the camera simultaneously captured four-channel reflected images. The data records are shown in Table 1:

[0068] Table 1 Step 2 Experimental Record

[0069]

[0070] Experiments show that step 2 effectively overcomes the complex interference on the coating surface through asymmetric Gabor filtering and dynamic noise suppression mechanism, improves detection accuracy and efficiency, and meets industrial-grade high robustness requirements.

[0071] In step 3 of the above embodiment, the working principle of step 3 is:

[0072] The binary scratch candidate area is input into the Gaussian pyramid decomposition module, and the scale factor is The recursive downsampling generates a 5-layer pyramid image {L0, L1, ..., L4}, and each layer is downsampled after convolution with a dynamic Gaussian kernel function. The dynamic Gaussian kernel function expression is:

[0073]

[0074] In formula (5), x and y are the two-dimensional coordinates of the image pixel points; δ is the standard deviation of the Gaussian kernel, which is used to control the blur intensity. The δ increases exponentially with the level and is defined as δ i =2 i / 2 (i=0,1,2,3,4); where δ i Represents the standard deviation of the Gaussian kernel of the i-th Gaussian pyramid; i is the pyramid level index;

[0075] Perform morphological opening operations on each layer of the pyramid to extract multi-scale features and define the feature map of the i-th layer The structural elements The radius r i =3×2 i ;L i Represents the pyramid image of the i-th layer; constructs the low-frequency layer L r With high frequency layer L j Attention weight map A r,j , where j = r + 1, the attention weight map A r,j The formula expression is:

[0076]

[0077] In formula (6), ||F r (x,y)||2 is L r The L2 norm feature strength of the layer; For L j The gradient L1 norm of the layer; F r (x,y) is the feature map after the r-th layer morphological opening operation; is the gradient magnitude of the j-th layer image; k is the pyramid level traversal index;

[0078] The low-frequency continuous features and high-frequency detail features are fused through weighted residual connections to generate an enhanced feature map F. fused , the formula expression is:

[0079]

[0080] In formula (7), A i,i+1 Represents the feature correlation weight between the low-frequency layer i and the high-frequency layer i+1, which is used to dynamically control the contribution of features at different levels in the fusion process; ⊙ represents element-by-element multiplication; ↑ represents the bilinear interpolation upsampling operation, which is used to convert the low-resolution feature map F i Upsample to high resolution; F i+1 Represents the high-frequency feature map of the i+1th layer;

[0081] Based on the polarized reflection image sequence in step 1, calculate the theoretical value of the refraction angle θ of the scratch area theory , the calculation formula is:

[0082]

[0083] In formula (8), n air and n glass are the refractive indices of air and glass, θ inc is the incident angle of the polarized light source in step 1; extract the enhanced feature map F fused The gradient direction θ of the edge point edge , if ||θ theory -θ edge If || is greater than the preset angle deviation threshold, it is determined to be an artifact and removed, and the fine screening scratch area map θ is output. refined .

[0084] Among them, the Gaussian pyramid decomposition module uses recursive downsampling to construct a 5-layer image space, and its dynamic Gaussian kernel function achieves scale adaptability through an exponentially increasing strategy of the standard deviation σ_i: i =3×2 i The exponential growth model of the layer i∈[0,4] makes the low layer (i=0) use a fine scale (r0=1.2) to retain high-frequency details, and the high layer (i=4) uses a coarse scale (r4≈2.49) to enhance low-frequency continuity, which is consistent with the scale space generation law of visual perception. The morphological opening operation is performed through the radius r i The disk structure element performs topological filtering on each layer of the pyramid, and its expansion-erosion operation can suppress isolated noise points and maintain scratch connectivity. iThe exponential growth of matches the scale expansion of the Gaussian kernel, forming a synergistic effect of multi-level morphological closure operations. The cross-layer attention mechanism introduces the normalized weight calculation of the feature intensity L2 norm and the gradient L1 norm, where the feature intensity L2 norm reflects the energy distribution of the regional continuity feature, and the gradient L1 norm represents the directional entropy of the edge sharpness. The two realize the dynamic coupling of low-frequency and high-frequency features through the soft attention function of formula (6). Its physical meaning is to enhance the spatial coherence and directional consistency of the scratch area. The weighted residual connection uses bilinear interpolation upsampling to solve the resolution mismatch problem of multi-scale features, and realizes the spatial modulation of high-frequency features through element-by-element multiplication. At the same time, the correlation weight λ_i is introduced to adjust the contribution between levels. This parameter is determined by the eigenvalue decomposition of the covariance matrix to ensure the geometric consistency of cross-layer feature fusion. The refraction angle consistency verification is based on the Fresnel reflection law. The theoretical refraction angle θ is calculated by formula (8) theory , which is essentially a deformed expression of Snell's law under the polarization reflection model, combined with the gradient direction angle θ edge The cosine similarity criterion can eliminate the pseudo edge response caused by coating interference, where the angle deviation threshold is set according to the angular resolution limit of the optical system under the Rayleigh criterion.

[0085] During implementation, dynamic Gaussian kernel convolution and downsampling are performed in parallel by CUDA kernel functions, and the initial δ i =1.0, the level increment coefficient is hard-coded to 2 i / 2 The morphological opening operation is realized by FPGA in real time, and the radius of the structural element r i =3×2 i Through the pre-storage of the lookup table, the erosion and expansion operations are completed through the shift register pipeline architecture. The cross-layer attention weight calculation module is integrated into the GPU. The L2 norm energy calculation adopts the fast square sum algorithm. The gradient amplitude is extracted by the Sobel operator. The normalized denominator is updated in real time by the global accumulator. In the residual fusion, the bilinear upsampling is accelerated by texture memory, and the weight matrix A is used. i,i+1 Stored as a 32-bit floating point array. The refraction angle verification module runs on the ARM Cortex-A72 processor, θ theory The polarized incident angle θ obtained in step 1 inc (Default is 60°) and the material refractive index (n air =1.0003,n glass =1.52) real-time calculation, edge gradient direction θ edgeExtraction is performed using a combined Canny operator and Hough transform, with an angle deviation threshold of 5° hard-coded into a logic comparator. Software parameter configuration includes 5 Gaussian pyramid levels, 2 morphological operation iterations, α for attention weight fusion, 1.2 for residual connection gradient compensation, and 3 pixels for dilation of the refraction angle verification mask. The processing flow is packaged in ROS, taking as input the binarized image and polarization parameters from step 2, and outputting a fine-screened scratch area map. Single-frame processing latency is ≤50ms (4K resolution), and peak GPU memory usage is ≤4GB.

[0086] The hardware for step 3 includes: NVIDIA A100 GPU (parallel computing), Xilinx Zynq UltraScale+ FPGA (real-time processing), 20-megapixel global shutter CMOS industrial camera (image input), multi-angle polarized light source array (ring layout), PCIe 3.0 bus (data transmission), and DDR4 memory (data cache). The GPU and FPGA are interconnected via PCIe. The light source array is arranged in a ring with a 30° inclination around the inspection station, and the camera lens is perpendicular to the coated lens surface. Based on the above hardware, a comparative experiment was designed to verify the performance advantages of step 3 in multi-scale feature fusion and artifact suppression. The control group (Group B) used a single-scale morphological opening operation (fixed 5×5 disk kernel) combined with traditional threshold segmentation. The experiment used 10 coated glass lenses (scratch width 2-10μm), and each group was tested 5 times. The experimental data are shown in Table 2:

[0087] Table 2 Step 3 Experimental Record

[0088]

[0089] Experiments show that step 3 effectively improves the integrity and artifact suppression capabilities of multi-scale feature fusion through dynamic Gaussian kernel decomposition, cross-layer attention weighting, and refraction angle verification, overcoming the problems of missed detection and artifact misjudgment of intermittent scratches by traditional methods, and meeting the high-precision requirements of complex industrial scenarios. Therefore, this step effectively suppresses artifact interference on the coating surface through the synergistic effect of the cross-layer attention weighting mechanism and physical refraction angle verification, preserving the microscopic details and macroscopic continuity of tilted and intermittent scratches. Dynamic Gaussian kernel decomposition balances noise suppression and feature retention, residual fusion enhances the complementarity of multi-scale information, and refraction angle constraints eliminate optical inconsistency artifacts, significantly improving detection accuracy and robustness in complex industrial scenarios and providing high-confidence input for subsequent 3D modeling.

[0090] In step 4 of the above embodiment, the direction-constrained convolution kernel group extracts the gradient direction histogram of each pixel point based on the fine-screened scratch area map, generates the scratch main direction distribution matrix through Hough transform, and constructs a multi-channel direction-constrained mask in combination with the refraction angle consistency direction encoding matrix output in step 3. Subsequently, the traditional isotropic convolution kernel is reconstructed into a direction-sensitive kernel structure using a differentiable direction parameterization method, and the mask matrix is mapped to the convolution kernel space through the kernel weight direction encoding function to generate a dynamic convolution kernel group within the restricted direction interval, wherein each kernel unit dynamically adjusts the weight distribution within the kernel according to the local direction field. When the direction-constrained convolution kernel group performs the restricted direction convolution operation, a sliding window is used to traverse the fine-screened area map. If the direction encoding value of the window center point exceeds the preset direction interval threshold, the kernel weight zeroing mechanism is triggered to mask the invalid direction response. Finally, the convolution output is texture decomposed using a multi-scale Gabor filter group, and the high-frequency scratch texture is separated from the low-frequency background baseband based on the wavelet packet transform. The multi-scale feature map is spatially projected using the GS orthogonalization algorithm to output the decoupled orthogonal encoded feature vector.

[0091] Further, if Figure 5 As shown, the multi-scale texture decomposition module includes a dual-tree complex wavelet decomposition unit, a direction-sensitive filtering unit, a cross-scale feature fusion unit, a texture orthogonal quantization unit, a non-local similarity aggregation unit and a spatial orthogonal coding generation unit; the dual-tree complex wavelet decomposition unit is used to decompose the input image into multi-scale high-frequency sub-bands and low-frequency approximate components through dual-tree complex wavelet transform, the high-frequency sub-bands cover six directions of ±15°, ±45° and ±75°, and the low-frequency approximate components retain the global texture contour; the direction-sensitive filtering unit is used to extract the main direction information based on the direction coding matrix DCM, and select the high-frequency sub-band that matches the scratch direction for direction enhancement filtering; the cross-scale feature fusion unit is used to use the cross-scale The attention mechanism calculates the correlation between adjacent scale sub-bands. If the correlation coefficient is higher than a preset threshold, weighted fusion is performed; otherwise, the independent scale features are retained. The texture orthogonal quantization unit is used to generate an energy quantization mask by calculating the ratio of the high-frequency sub-band energy to the low-frequency component. If the high-frequency energy ratio of any area exceeds the preset threshold, it is marked as a scratch candidate area. The non-local similarity aggregation unit is used to perform non-local mean filtering in the area marked by the energy mask, search for similar texture blocks through a block matching algorithm, and perform weighted averaging. The spatial orthogonal coding generation unit is used to perform principal component analysis on the denoised texture map, extract the principal component vector, and splice it with the directional coding matrix DCM to generate an orthogonal coding vector.

[0092] Based on the gradient direction histogram of the finely screened scratch region image, the Hough Transform is used to generate the main direction distribution matrix. Its mathematical essence is to map straight line features through parameter space. The Hough Transform detects the directional distribution of straight line segments in the image by constructing an accumulator matrix in the parameter space (ρ, θ). Combined with the refraction angle consistency direction encoding matrix output in step 3, a multi-channel direction constraint mask is constructed. Its physical significance lies in constraining the theoretical range of scratch directions through the optical refraction law (Snell's law). This fusion of the geometric direction detection of the Hough Transform with the physical refraction characteristics forms a multimodal encoding of the direction field.

[0093] Traditional isotropic convolution kernels (such as Gaussian kernels) are reconstructed into direction-sensitive kernel structures through differentiable directional parameterization methods. Specifically, the kernel weight direction encoding function maps the mask matrix to the kernel space. In essence, it introduces a directional weight modulation factor so that the convolution kernel exhibits direction selectivity in the frequency domain. For example, dynamic adjustment of the weight distribution within the kernel is achieved through the directional sensitivity of the Gabor kernel (frequency and directional parameter coupling) or the channel weight adjustment of the depthwise separable convolution. This reconstruction utilizes the geometric transformation invariance of the kernel function, converting the directional constraint into a weight bias in the kernel space, ensuring that the convolution operation activates the response only within the preset directional range.

[0094] During the convolution operation, a sliding window is used to traverse the fine-screening region map. When the directional encoding value at the center of the window exceeds a preset threshold, the kernel weight reset mechanism is triggered. This mechanism is implemented by the element-by-element product of the directional mask matrix and the convolution kernel. Mathematically, it can be expressed as: Kactive = K⊙Mdirection, where Mdirection is a binary mask matrix that takes on a value of 1 only within the valid directional range. This dynamic masking strategy effectively suppresses background noise and false directional interference, improving the signal purity of scratch features.

[0095] The Gabor filter bank achieves texture decomposition through multi-scale frequency and directional selectivity. Its mathematical form is the product of a complex Gaussian envelope modulated sine wave, expressed as: G(x,y) = ex′² + γ²y′² 2σ² · ei(2πfx′ + φ), where x′ = xcosθ + ysinθ and y′ = -xsinθ + ycosθ, with θ being the directional parameter. By setting high-frequency subbands at ±15°, ±45°, and ±75°, the possible main directions of scratches are covered, while retaining the low-frequency approximation to maintain the global texture profile.

[0096] The dual-tree complex wavelet transform generates complex coefficients through two independent filter trees, achieving near-translation invariance and multi-directional selectivity (six directions). Its physical significance lies in decomposing the image into high-frequency details (scratch texture) and a low-frequency baseband (background contour). A direction-sensitive filter unit selects the high-frequency subband that matches the scratch direction for enhancement, for example, by applying nonlinear amplification to specific subbands using a directional weight matrix.

[0097] The cross-scale attention mechanism dynamically adjusts the fusion weights by calculating the correlation coefficient between adjacent scale subbands. If the correlation coefficient is above a threshold (e.g., 0.7), weighted fusion is performed to enhance continuous scratch features; otherwise, independent scale features are retained to maintain local details. The texture orthogonal quantization unit generates an energy quantization mask based on the ratio of high-frequency subband energy to low-frequency components. The mathematical expression is: Eratio = ∑|Hhigh|2∑|Llow|2. When Eratio > Tthreshold, the region is marked as a scratch candidate. This method achieves region segmentation based on the fact that the high-frequency energy of the scratch region is significantly higher than that of the background.

[0098] Non-local mean filtering uses a block matching algorithm to search for similar texture blocks and employs weighted averaging to suppress random noise. Its core lies in constructing a non-local similarity metric function: w(p,q) = exp(-||Bp-Bq||2h²), where Bp and Bq represent image blocks and h is a smoothing parameter. Principal component analysis (PCA) performs dimensionality reduction on the denoised texture image, extracting the principal component vector in the direction of maximum variance. This is then concatenated with the directional coding matrix (DCM) to generate an orthogonal coding vector. The Gram-Schmidt orthogonalization algorithm is used to spatially project the eigenvectors, ensuring independence between dimensions and forming a decoupled orthogonal code.

[0099] In one implementation, the system used a FLIR Blackfly S BFS-PGE-51S5P-C four-channel polarization camera as the primary imaging unit, paired with a Mitutoyo M Plan Apo 50× aspheric microscope objective lens with a 13mm working distance and an imaging resolution of 2.45μm / pixel. The camera was connected to the camera via a C-Mount interface, ensuring high-resolution image acquisition. The camera was mounted directly above a six-axis translation stage (PI H-824, with ±0.1μm repeatability). The stage communicated via EtherCAT with a NIPCIe-7842 motion control card in an industrial computer, enabling XY scanning of the sample and dynamic focus compensation in the Z axis. The imaging light source was a Thorlabs LIU630A annular polarized LED, coaxially mounted at a 45° angle around the camera. The light source had a wavelength range of 450-650nm and a polarization degree >95%. The light source controller controlled the brightness and triggered the exposure synchronization signal via an RS-485 interface. To ensure imaging stability, the system integrates the Keyence LK-G5000 laser ranging module, which provides real-time feedback of the distance data between the objective lens and the sample via Gigabit Ethernet, dynamically calibrates the focusing parameters, and controls the depth of field error to <0.5μm. The computing core adopts the collaborative architecture of NVIDIA A6000GPU (48GB video memory) and Xilinx XC7K325T FPGA. The GPU is responsible for multi-scale feature fusion and texture decomposition. The FPGA is directly connected to the GPU via PCIe Switch (bandwidth 64GB / s) to achieve hardware acceleration of the direction-constrained convolution kernel group. In terms of hardware electrical connection, the camera, light source, displacement platform and computing unit are connected through a hybrid network of PCIe 3.0, EtherCAT and RS-485. The power supply system uses an independent 24V DC power supply and uses an optoelectronic isolator to suppress electromagnetic interference to ensure the stability and real-time performance of signal transmission.

[0100] Step 4 is based on physical optical properties and deep learning acceleration technology, and real-time processing is achieved through parameterized modeling and parallel computing. The construction of the direction-constrained convolution kernel group takes the fine-screened scratch area map (512×512 binary image) and the refraction angle encoding matrix (θ∈[0°,180°], step size 1°) as input, and uses the improved Hough transform algorithm (ρ step size 0.5 pixels, θ step size 0.5°) to generate the main direction distribution matrix DCM, and combines the differentiable direction parameterization method to reconstruct the 5×5 convolution kernel. The weight function is w(θ)=cos2(θ-θ0), where θ0 is the local main direction, and the direction deviation exceeds the threshold θ t=5°, triggering the FPGA logic gate to reset the kernel weights to zero. The multiscale texture decomposition module uses the Dual-Tree Complex Wavelet Transform (DTCWT) algorithm with a decomposition level of L = 5, covering six high-frequency directional subbands at ±15°, ±45°, and ±75°. Low-frequency components retain the global contour. The Matlab code is implemented using dtcwt2(img,'Level',5,'qshift','qshift_06'). Non-local means filtering uses a Bayesian optimization algorithm with a block size of 7×7, a search area of 21×21, a smoothing parameter h = 0.3σ (σ is the noise standard deviation), and a weighting function w(p,q) = exp(-||Bp-Bq||2 / h2). This is accelerated using CUDA kernels. In the orthogonal code generation stage, principal component analysis (PCA) reduces the feature dimensionality to 3D, and then combines the Gram-Schmidt algorithm to project it into orthogonal space. Matrix operations are accelerated using the CuBLAS library. In the real-time processing pipeline design, the polarization camera triggers exposure at a 60fps frame rate. After the image is downsampled to 512×512 by the GPU, the FPGA completes the direction-constrained convolution within 5ms, and the GPU completes texture decomposition and encoding within 15ms. The end-to-end latency is less than 30ms, meeting the real-time requirements of the production line. Key parameters include the Hough transform accumulator threshold (30% of the maximum number of votes), the Gabor filter (λ = 10 pixels, b = 1), and the cross-scale fusion correlation coefficient threshold (0.7).

[0101] The system achieves highly robust detection through multimodal signal fusion and coordinated debugging of hardware and software. In terms of hardware integration, the signal flow adopts a layered topology: polarization camera data is transmitted directly to the GPU / FPGA via a PCIe switch, motion control commands are sent to the displacement platform via the EtherCAT bus, and light source trigger signals are synchronized via RS-485 to ensure timing consistency. The power system provides independent grounding for the camera, light source, and FPGA, with a common ground impedance of <0.1Ω to prevent electromagnetic interference. At the software level, detection results are optimized through parameter tuning and performance testing. For example, the Hough transform accumulator threshold is set to 30% of the maximum number of votes, the Gabor filter orientations θ = 0°, 45°, and 90°, and the non-local mean filter block matching algorithm achieves 10^6 similarity calculations per second on the NVIDIA A6000.

[0102] In this implementation, a comparative experiment was conducted to verify the performance advantages of step 4 (Group A) over traditional isotropic convolution (σ=1.5) + fixed threshold segmentation (Group B) in terms of directional sensitivity and background decoupling. The experimental data is shown in Table 3:

[0103] Table 3 Step 4 Experimental Record

[0104]

[0105] As shown in the table, Group A achieved an average detection accuracy of 96.1% (79.8% for Group B), a false positive rate of 2.3% (15.0% for Group B), and a single-frame processing time reduction of approximately 22% (55.9ms vs. 71.5ms). The directional consistency score (0.92 vs. 0.59), background decoupling ratio (27.1dB vs. 12.9dB), and feature orthogonality (0.88 vs. 0.53) were all significantly superior to Group B. Experiments have shown that Step 4, through directional constrained convolution and multi-scale orthogonal decomposition, effectively improves the directional detection capability and background suppression of scratches on complex surfaces, overcoming the directional confusion and noise coupling issues of traditional methods and meeting the requirements of industrial-grade high-precision detection. Therefore, compared with traditional methods, this technology significantly improves the detection accuracy of sub-pixel scratches in complex coating environments through the collaborative optimization of dynamic parameterized reconstruction of direction-constrained convolution kernel groups and multi-scale frequency domain orthogonal decomposition: the Hough transform-Legendre polynomial mapping mechanism of the direction-sensitive kernel effectively suppresses non-target directional interference, the non-redundant decomposition characteristics of the dual-tree complex wavelet transform and the cross-scale fusion strategy enhance the high-frequency texture representation capability, the non-local similarity aggregation and PCA-GS orthogonalization algorithm jointly eliminate background noise and feature redundancy, and the embedding of physical constraints on refraction angles further reduces the false detection rate of artifacts, achieving collaborative optimization of scratch detection sensitivity and anti-interference performance in industrial high-noise and low-contrast scenes, meeting the high real-time and high robustness detection requirements of production lines.

[0106] In step 5 of the above embodiment, if Figure 6As shown, the cascade network model includes: a multi-scale feature pyramid input layer, a direction distribution thermal estimation layer, a dynamic FP16 quantization allocation layer, a three-dimensional deformable convolution layer, a residual attention fine inspection layer and a geometric parameter regression layer; the multi-scale feature pyramid input layer is used to generate a multi-scale feature pyramid based on the decoupled feature map input through strided convolution and bilinear interpolation, fuse low-resolution semantic information with high-resolution detail features, and output a pyramid feature tensor; the direction distribution thermal estimation layer is used to perform 3×3 direction-sensitive convolution on the pyramid feature tensor to generate a direction distribution thermal map, and calculate the probability distribution of the scratch direction of each pixel through the Softmax function, and output a thermal confidence matrix; the dynamic FP16 quantization allocation layer is used to trigger FP when the regional confidence of the thermal confidence matrix is greater than or equal to 0.8. 32 full-precision calculation; when the regional confidence is greater than or equal to 0.5 and less than 0.8, FP16 mixed-precision calculation is enabled; when the regional confidence is less than 0.5, the calculation is skipped, and the quantization mask and the dynamic calculation feature map are output; the three-dimensional deformable convolution layer is used to extract the three-dimensional geometric deformation features of the dynamic calculation feature map through the deformable convolution kernel, and the kernel offset of the deformable convolution kernel is dynamically adjusted by the scratch direction probability to generate a three-dimensional deformation feature tensor; the residual attention fine inspection layer is used to fuse the coarse screening branch and the fine inspection branch features of the three-dimensional deformation feature tensor through the residual connection, and output the fine inspection feature map; the geometric parameter regression layer is used to regress the scratch length, width, depth and curvature parameters of the fine inspection feature map through the fully connected layer and the three-dimensional coordinate mapping, and output a topological structure encoding matrix containing geometric parameters and depth information.

[0107] Among them, the multi-scale feature pyramid input layer uses strided convolution (step size 2) and bilinear interpolation (scaling factor 0.5) to generate P2-P5 multi-resolution feature maps, where P2 retains high-resolution details (such as microcracks) and P5 encodes global semantic information (such as scratch direction). Cross-scale feature fusion is achieved through channel splicing and 1×1 convolution, enhancing the model's sensitivity to tiny defects and its ability to capture macroscopic morphology.

[0108] The direction distribution thermal estimation layer extracts the azimuth response of the feature map through a 3×3 direction-sensitive convolution kernel (8 directions, 45° interval) to generate a direction distribution heat map. The direction probability of each pixel is normalized to the interval [0, 1] by the Softmax function. The thermal confidence matrix quantifies the regional direction consistency through the direction entropy (low entropy value indicates direction concentration).

[0109] The dynamic FP16 quantization allocation layer dynamically switches the calculation precision according to the thermal confidence: the high confidence area (≥0.8) uses FP32 full precision to retain geometric details, the medium confidence area (0.5~0.8) enables FP16 mixed precision (scaling factor 256) to compress the data width, and the low confidence area (<0.5) skips the calculation to reduce redundancy. The quantization mask is generated in real time through bit operation instructions.

[0110] The 3D deformable convolution layer learns the local geometric deformation offsets Δx and Δy through the deformable convolution kernel (DCNv2). The offsets are linearly mapped by the directional probability matrix (Δ=α·P_dir+β). The weights in the kernel are dynamically adjusted with the offsets to enhance the 3D deformation modeling capability of curved scratches.

[0111] The residual attention fine inspection layer adopts the channel-spatial dual attention mechanism (CBAM). The channel attention generates weights through global average pooling and fully connected layers, the spatial attention generates masks through 7×7 convolution, and the residual connection fuses the coarse screening and fine inspection branch features to suppress background interference.

[0112] The geometric parameter regression layer regresses the scratch length (based on arc length integral), width (lateral gradient extreme value difference), depth (reflection intensity-depth mapping model) and curvature (second-order derivative fitting) through a fully connected layer (dimension 512→128) and three-dimensional coordinate mapping (homogeneous matrix transformation), and outputs a topological structure encoding matrix.

[0113] The cascaded network model is implemented on an NVIDIA A100 GPU. The decoupled feature maps are fed into the GPU memory via the PCIe 4.0x16 bus, and multi-scale pyramid construction is performed in parallel using CUDA kernels (thread block size 32×32). The parameters of the direction-sensitive convolution kernels are stored in texture memory, and the softmax calculation is accelerated using the cuDNN library. Dynamic FP16 quantization is implemented using the TensorRT engine, with FP32 / FP16 switching thresholds set to 0.8 / 0.5, and the quantization mask updated in real time via bitmask operations. The 3D deformable convolution layer is based on the PyTorch DCNv2 module, with linear offset mapping coefficients α = 0.5 and β = -0.1, and a kernel size of 3×3. The residual attention module is optimized using the TVM compiler, and the channel attention fully connected layer is reduced to 64 dimensions to reduce computational overhead. The geometric parameter regression layer is accelerated using ONNX Runtime, with fully connected layer parameters preloaded into GPU constant memory. 3D coordinate mapping is implemented using OpenGL compute shaders. In the software parameter configuration, the pyramid levels are P2-P5, the number of directions is 8, the FP16 scaling factor is 256, the deformable convolution learning rate is 0.01, the residual attention expansion coefficient is 2, the processing frame rate is 25fps (4K resolution), the GPU memory usage peak is 12GB, and the end-to-end latency is ≤80ms.

[0114] Compared with traditional static networks and fixed-precision calculations, this step significantly improves the geometric parameter estimation accuracy of complex scratches while reducing computing resource consumption through the collaborative optimization of dynamic precision allocation and three-dimensional deformation modeling; the directional thermal-driven quantization strategy balances detection efficiency and detail retention capabilities, and the residual attention mechanism enhances the robustness of low-contrast defect discrimination, providing high-precision, low-latency scratch three-dimensional topology analysis capabilities for high-speed production lines.

[0115] In step 6 of the above embodiment, the working principle of step 6 is:

[0116] The coating process parameter data stream is read in real time through the CAT bus protocol. A multi-physics field coupling model is constructed based on the film thickness and refractive index. The minimum wavelength offset algorithm is used to calculate the compensation light source wavelength adjustment amount Δλ. The calculation formula is:

[0117]

[0118] In formula (9), λ0 is the initial compensation light source; n film is the refractive index of the coating layer; n air is the refractive index of air; d is the thickness of the coating layer;

[0119] Trigger the near-infrared compensation light source in step 1 to switch to λ new =λ0+Δλ; Project the scratch topology output in step 5 to the standard template coordinate system, and extract the offset error matrix E between the actual coordinates and the theoretical coordinates through the Lissajous pattern registration algorithm proj , and generates the incident angle correction Δθ based on the error gradient field inc , the formula expression is:

[0120]

[0121] In formula (10), ||E proj || F Represents the projection error matrix E proj The F norm of the system is used to quantify the overall projection deviation of the system; D base is the reference distance of the standard sample;

[0122] Dynamically adjust the incident angle of the polarized light source in step 1 to θ′ inc =θ inc +Δθ inc ; At the same time, the frequency domain characteristics of the error matrix are analyzed, and the convolution kernel offset optimization function is used to calculate the deformation vector δ of the three-dimensional deformable convolution kernel in step 4 k , the calculation formula is:

[0123]

[0124] In formula (11), Represents the projection error to the convolution kernel weight W k The partial derivative of is used to reflect the sensitivity of the weight to the error; η is the learning rate coefficient;

[0125] Update the convolution kernel weight W′ k =W k +δ k ; and fuse Δθ through the Bayesian incremental learning algorithm inc and δ k Go to the optical collaborative parameter library to generate a versioned parameter set.

[0126] Among them, the CAT bus protocol (Controller Area Network Tunneling) uses an industrial Ethernet architecture (such as EtherCAT) to achieve high-speed data exchange between coating equipment and detection systems. Its physical layer uses differential signal transmission, and the protocol stack is streamlined to the physical layer, data link layer, and application layer, supporting multi-node synchronization under a master-slave architecture. By reading the coating process parameters (film thickness d, refractive index n) in real time, a multi-physics field coupling model is constructed. This model integrates the interaction of the thermal field (temperature gradient), the seepage field (film uniformity), and the stress field (film deformation) to derive the equivalent refractive index n_{eq} of the film's optical properties. Its mathematical form can be expressed as an extension of the Maxwell-Garnett equivalent medium theory. Based on this, the minimum wavelength shift algorithm dynamically adjusts the compensation light source wavelength λ_c = λ_0 + Δλ by solving the extreme value of the optical path difference Δ = 2d(n_{film} - n_{air}) to match the film's achromatic condition and eliminate interference noise caused by film thickness fluctuations.

[0127] Secondly, the Lissajous pattern registration algorithm uses the principle of orthogonal signal synthesis to project the scratch topology output in step 5 into the standard template coordinate system. The phase difference Δφ between the actual coordinates and the theoretical coordinates is extracted through Fourier transform, and the error matrix E∈R^{m×n} is constructed. Its F norm ||E proj || F Characterizes the global projection deviation, and its physical meaning is the root mean square error of the coordinate offset in Euclidean space. Error gradient field Generate the incident angle correction value through Jacobian matrix decomposition Where η is the learning rate coefficient, d_0 is the reference distance of the standard sample, and the correction amount adjusts the incident angle θ_c = 0_0 + Δθ through the principle of polarized light interference to compensate for the imaging distortion caused by the deviation of the light source incident angle.

[0128] Finally, the convolution kernel offset optimization function is based on the frequency domain characteristics of the error matrix and uses the back propagation algorithm to calculate the deformation vector of the three-dimensional deformable convolution kernel. in is the partial derivative of the error with respect to the kernel weight, reflecting the weight's sensitivity to projection bias. Using a Bayesian incremental learning algorithm, ΔW is fused with the historical parameter library through a Gaussian process to generate a versioned parameter set, enabling adaptive updates of the kernel weights W_c = W_0 + ΔW. Its mathematical foundation is maximizing the posterior probability within the expectation-maximization (EM) framework.

[0129] During implementation, the system uses an EtherCAT master (such as a Beckhoff CX2040 PLC) and coating equipment slaves (film thickness sensors and refractive index monitoring modules) in a daisy-chain topology via twisted-pair cables. The communication cycle is set to 1ms, and the data frame format follows the CoE (CANopen over EtherCAT) protocol. The system transmits film thickness d (accuracy ±0.1nm) and refractive index n (accuracy ±0.001) in real time. The compensation light source module uses a tunable near-infrared laser (wavelength range 1500-1600nm). A PID controller adjusts the piezoelectric displacement of the Bragg grating to achieve subnanometer adjustment of λ_c (resolution 0.01nm).

[0130] In the software implementation, the multi-physics coupling model is solved using the finite element method (FEM), the mesh is divided into unstructured tetrahedrons, the boundary condition is set to the thermal radiation coefficient ε of the coating cavity = 0.85, and the calculation step size Δt = 10ms. Lissajous figure registration is implemented through the FPGA (Xilinx XC7K325T) to implement the fast Fourier transform (FFT), parallelize the calculation of the frequency domain components of the error matrix E, and solve the phase difference Δφ through the CORDIC algorithm, with a processing delay of <2ms. The convolution kernel offset optimization runs on the NVIDIAA100 GPU and is accelerated by the CUDA core. The learning rate η is 0.01, and the weight update cycle is synchronized with the coating process (60Hz). The optical collaborative parameter library is versioned through a PostgreSQL database, using time series partitioning to store historical parameter sets and supporting Markov Chain Monte Carlo (MCMC) sampling for Bayesian incremental learning.

[0131] In another implementation, an online monitoring system for optical lens transmittance includes: a multimodal optical acquisition module, a gradient domain dynamic processing module, a multi-scale feature fusion module, a directional constraint decoupling module, a three-dimensional topology modeling module, and a dynamic closed-loop optimization module;

[0132] The multimodal optical acquisition module is used to integrate a four-directional polarized light source array and a near-infrared compensation light source, generate a composite light field with orthogonal polarization states through time-sharing triggering, synchronously project the compensation light spot onto the coating surface, and output a four-channel polarized reflection image sequence and a coating interference compensation map;

[0133] The gradient domain dynamic processing module is used to perform multi-directional gradient convolution on the four-channel polarized reflectance image sequence based on the asymmetric Gabor filter bank, dynamically optimize the segmentation threshold of the improved OTSU algorithm in combination with the noise distribution characteristics of the coating interference compensation image, and output the binary scratch candidate area and the scratch direction encoding matrix;

[0134] The multi-scale feature fusion module is used to input the binary scratch candidate area into the Gaussian pyramid decomposition network, extract multi-scale morphological features in layers, fuse low-frequency continuous features with high-frequency detail features through a cross-layer attention weighting mechanism, and eliminate artifact interference based on the refraction angle consistency verification algorithm, and output a high-confidence scratch area map after fine screening;

[0135] The directional constraint decoupling module is used to construct a directional constraint dynamic convolution kernel group based on the directional coding matrix and the finely screened scratch area map, perform multi-scale texture orthogonal decomposition within the restricted directional range, quantize and separate scratch features and base texture noise, and output a decoupled feature map and a spatial orthogonalized code;

[0136] The three-dimensional topology modeling module is used to input the decoupled feature map into the cascade network model, dynamically allocate FP16 quantization computing resources based on the directional distribution heat map, and fuse the spatial orthogonalization coding to generate a scratch topology structure containing three-dimensional depth information and geometric parameters;

[0137] The dynamic closed-loop optimization module is used to obtain the coating thickness and refractive index parameters in real time through the CAT bus, perform back-projection error analysis on the scratch topology structure in combination with a standard template, dynamically adjust the compensation light source wavelength of the multimodal optical acquisition module and the convolution kernel offset of the directional constraint decoupling module, and update the light field collaborative parameter library.

[0138] When implementing, Figure 2 As shown, the four-channel polarized reflectance image sequence (RGB + near-infrared) from the multimodal optical acquisition module is transmitted via a Camera Link interface to the FPGA (Xilinx Zynq UltraScale+) in the gradient domain dynamic processing module. The coating interference compensation map is then transferred in parallel to the GPU memory (NVIDIA A100) via a PCIe 4.0x8 bus. The FPGA's GPIO pins control the time-sharing triggering of the light source array via LVDS signals, ensuring synchronization of exposure timing with image acquisition (error ≤ 1μs). The polarization image sequence is streamed in 12-bit RAW format with a bandwidth of 4.8Gbps; the coating compensation map is stored in 32-bit floating-point format in a DDR4 memory pool.

[0139] The gradient domain dynamic processing module binarizes the scratch candidate region map (binary bitmap) and the direction encoding matrix (32-bit floating-point tensor) and transmits them to the GPU cluster via NVLink 3.0 high-speed interconnect for invocation by the multi-scale feature fusion module. The asymmetric Gabor kernel parameters are stored in the FPGA's block RAM, with a kernel size of 9×9 and 12 directions. The dynamic bandwidth adjustment parameter (σ=1.2) is updated in real time via the SPI interface. The segmentation threshold dynamic table is shared with the GPU's global memory via PCIe memory mapping (MMIO) for invocation by the improved OTSU algorithm.

[0140] The high-confidence scratch region map (binary mask) obtained through the multi-scale feature fusion module is transferred via a PCIe switch (PLX PEX8796) to the FPGA (Xilinx Alveo U280) in the direction constraint decoupling module. The multi-scale feature map is also shared via a DDR4 memory pool (Micron 3200MHz). Cross-layer attention weights are cached in the GPU's L2 cache, and the weight matrix is updated in real time via atomic operations (Atomic Add). The refraction angle consistency verification results are written to the FPGA's configuration registers in 8-bit mask format via a DMA channel.

[0141] The decoupled feature map (128-dimensional floating-point tensor) and spatial orthogonalized code (256-dimensional vector) from the directional constraint decoupling module are transmitted to the Jetson AGX Xavier in the 3D topology modeling module via a QSFP28 optical module (100Gbps). The kernel weights of the directional constraint convolution kernel group are burned into on-chip BRAM via the FPGA's JTAG interface, and the kernel offsets (Δx, Δy) are updated in real time via PCIe P2P DMA. The orthogonalized code stream is transmitted using the RoCEv2 (RDMA over Converged Ethernet) protocol to reduce CPU load.

[0142] The scratch topology (including 3D coordinates, curvature, depth, and other parameters) from the 3D topology modeling module is transmitted via EtherCAT (Beckhoff EK1100) to the industrial PLC (Siemens S7-1500) of the dynamic closed-loop optimization module. The backprojection error matrix is transmitted back via PCIe Gen4 to the FPGA of the multimodal optical acquisition module, driving light source wavelength adjustment (Δλ) and incident angle calibration (Δθ). The deformation vector δ_k, which updates the convolution kernel offset, is written in real time to the FPGA of the direction constraint decoupling module via the AXI-Stream interface.

[0143] The dynamic closed-loop optimization module is connected to the multimodal optical acquisition module and the direction constraint decoupling module; the film thickness (d) and refractive index (n_film) are read in real time via the EtherCAT bus, and the light source wavelength adjustment instruction is sent via I 2 The optimized convolution kernel weights are written directly into the FPGA configuration space via PCIe P2P (Peer-to-Peer), bypassing CPU intervention.

[0144] During implementation, the system's acquisition side: the light source array and industrial camera achieve sub-microsecond timing alignment via a fiber-optic synchronization link (PTP 1588), ensuring spatiotemporal consistency of multimodal data. On the computing side: a GPU cluster and FPGA form a heterogeneous computing unit via NVLink / PCIe. The GPU processes intensive algorithms (such as feature fusion), while the FPGA is responsible for low-latency control (such as dynamic kernel adjustment). The polarization image stream (4K@60fps) is transmitted full-duplex via Camera Link, occupying independent PCIe channels to avoid bandwidth contention. Closed-loop calibration instructions are triggered by FPGA hardware interrupts (IRQs), with a response time of ≤10μs. The light source and computing unit are each powered by redundant DC power supplies (Meanwell HEP-1000) to prevent single points of failure. The transmission link has a built-in CRC-32 checksum, and erroneous data packets are retransmitted via the PCIe Retry mechanism.

[0145] Although specific embodiments of the present invention have been described above, those skilled in the art will appreciate that these specific embodiments are merely illustrative, and that those skilled in the art may omit, substitute, and modify the details of the methods and systems described above without departing from the principles and spirit of the present invention. For example, combining the above method steps to perform substantially the same functions in substantially the same manner to achieve substantially the same results falls within the scope of the present invention. Accordingly, the scope of the present invention is limited solely by the appended claims.

Claims

1. A method for detecting scratches on a glass lens surface, characterized in that: The following steps are involved: Step 1: Based on the multi-angle polarized light source array, a four-directional polarizer and a near-infrared compensation light source are integrated to generate a composite light field with orthogonal polarization states. The compensation light spot is synchronously projected onto the coating layer through a time-sharing trigger mechanism, and a four-channel polarized reflection image sequence and a coating interference compensation map are output; Step 2: Based on the four-channel polarized reflectance image sequence, a directional gradient tensor matrix is generated through an asymmetric Gabor filter bank. Combined with the noise distribution of the coating interference compensation image, the threshold segmentation parameters are dynamically adjusted through the improved OTSU threshold segmentation method to output the binary scratch candidate area and directional coding matrix; Step 3: Input the binary scratch candidate area into the Gaussian pyramid decomposition module, extract multi-scale morphological features layer by layer, fuse low-frequency continuous features with high-frequency detail features through a cross-layer attention weighting mechanism, use a refraction angle consistency verification algorithm to eliminate artifact interference, and output a fine-screened scratch area map; Step 4: Based on the directional coding matrix and the finely screened scratch area map, a directional constrained convolution kernel group is constructed, dynamic convolution operations are performed within the restricted directional interval, and scratches and background textures are quantized and decomposed through a multi-scale texture decomposition module, and a decoupled feature map and spatial orthogonalized code are output; Step 5: Input the decoupled feature map into the cascade network model, evaluate the confidence based on the directional distribution heat map, dynamically allocate FP16 quantization computing resources to high-probability areas, and output the scratch topology structure containing geometric parameters and 3D depth information; Step 6: Obtain coating process parameters, including film thickness and refractive index, in real time through the CAT bus, dynamically adjust the compensation light source wavelength of step 1, perform reverse projection error analysis on the output of step 5 through a standard sample, close-loop calibrate the light source incident angle of step 1 and the convolution kernel offset of step 4, and update the light synergy parameter library.

2. The method for detecting scratches on a glass lens surface according to claim 1, wherein: The working principle of step 1 is as follows: based on the wavelength polarization direction coding protocol, polarization angle and wavelength parameters are assigned to each LED unit in the multi-angle polarized light source array, wherein the visible light polarized light source group uses a four-way polarizer switching mechanism to output linearly polarized light with polarization angles of 0°, 45°, 90°, and 135° and wavelengths of 450nm and 650nm respectively; and synchronously triggers an 850nm near-infrared compensation light source, and projects the compensation light spot onto the interface between the coating layer and the substrate through a polarization beam splitter. The industrial camera captures a four-channel polarized reflection image sequence based on the synchronous trigger signal, and each channel is The reflected light intensity distribution of a single polarization angle and wavelength combination is obtained; then, the polarization component parallel to the incident plane and the polarization component perpendicular to the incident plane are separated by a four-channel polarization solution algorithm, and the polarized reflection image sequence of four time-sharing exposures is dynamically fused; in response to the interference of coating layer thickness fluctuations, the coating process parameters are obtained in real time through the CAT bus, and a closed-loop feedback control algorithm is used to drive the rotating motor to adjust the azimuth angle of the polarizer to the extinction position to reduce the mirror reflection light intensity; finally, the grayscale gradient covariance analysis method is used to calculate the pixel intensity variance matrix of the near-infrared compensation area and the non-compensation area to generate the coating interference compensation map.

3. The method for detecting scratches on a glass lens surface according to claim 1, wherein: The asymmetric Gabor filter bank is configured with differentiated aspect ratios and phase offset parameters along 12 directions. The ratio of the major axis to the minor axis of the horizontal and vertical kernels is 3:1, the ratio of the diagonal kernels is 2.5:1, and the ratio of the major axis to the minor axis in the medial direction is 2.8:

1. The working principle of the asymmetric Gabor filter bank is as follows: Through multi-directional convolution kernel, each pixel in the input polarized reflectance image sequence is parallelized to extract the gradient amplitude phase response characteristics of each pixel in the anisotropic frequency domain space, and generate 12-directional gradient feature vectors; During the initialization phase of the asymmetric Gabor filter bank, a dynamic bandwidth adjustment mechanism is used to adaptively adjust the half-peak bandwidth parameters of the kernel function in each direction according to the noise distribution characteristics of the coating interference compensation map; During the convolution process, the real and imaginary components of the orthogonal Gabor are asymmetrically weighted and fused using the directional weight factor, and the amplitude squared sum phase difference compensation algorithm is used to enhance the signal-to-noise ratio of the scratch edge directional response. Based on the 12 directional gradient feature vectors, a directional gradient tensor matrix that fuses directional coding and intensity distribution is constructed through a tensor aggregation method.

4. The method for detecting scratches on a glass lens surface according to claim 1, wherein: The working method of the improved OTSU threshold segmentation method is: S1, coating interference compensation diagram C map The noise distribution is input and the noise variance σ in the neighborhood of each pixel (x, y) is calculated. 2 (x,y), if σ is within the current window 2 (x, y) exceeds the preset noise variance threshold η, the window size adaptive reduction mechanism is triggered, and the window size is adjusted to the original size. Generate dynamic window W dyn ; S2, in the dynamic window W dyn In the , the weight factor α is introduced to optimize the inter-class variance to avoid over-segmentation in high noise areas. The optimized inter-class variance is defined as: In formula (1), α is the weight factor, which is defined as α = 1 / (1 + σ 2 ), used to balance global and local feature contributions; μ T is the global grayscale mean, which is used to preserve the image brightness distribution characteristics; μ0 and μ1 are the grayscale means of the foreground and background in the dynamic window respectively; ω0 and ω1 are the pixel weights of scratches and base respectively; S3, based on the coating interference compensation diagram C map Generate noise probability matrix N map (x,y), where N map (x,y)=C map (x,y) / max(C map ), define the noise suppression coefficient β=1-N map (x, y), the segmentation threshold of the high-probability noise area is shifted toward the background class by improving the segmentation decision function. The formula expression of the improved segmentation decision function is: S4, perform Gaussian pyramid downsampling on the directional gradient tensor matrix, and calculate the initial threshold T at the coarse scale layer of 1 / 4 resolution coarse , and update it layer by layer upwards. The formula is: In formula (3), λ represents the gradient compensation factor, which is defined as OGM fine Represents the directional gradient tensor matrix at a fine scale, which contains the gradient amplitude and direction information at a fine scale and is used to preserve the edge detail features of the scratch; S5. Calculate the directional discreteness θ of the pixels in the candidate area by combining the directional encoding matrix std , the calculation formula is: θ std =std(DCM(x,y)) (4) In formula (4), DCM(x,y) represents the direction encoding matrix of the pixel point (x,y), which is used to record the maximum gradient response direction of each pixel point; std represents the standard deviation operation, which is used to quantify the discrete degree of the pixel direction angle in the scratch candidate area; In θ std When it is greater than 15°, it is determined to be artifact interference and removed, and finally the binary scratch area and direction code are output.

5. The method for detecting scratches on a glass lens surface according to claim 1, wherein: The working principle of step 3 is: The binary scratch candidate area is input into the Gaussian pyramid decomposition module, and the scale factor is The recursive downsampling generates a 5-layer pyramid image {L0, L1, ..., L4}, and each layer is downsampled after convolution with a dynamic Gaussian kernel function. The dynamic Gaussian kernel function expression is: In formula (5), x and y are the two-dimensional coordinates of the image pixel points; δ is the standard deviation of the Gaussian kernel, which is used to control the blur intensity. The δ increases exponentially with the level and is defined as δ i =2 i / 2 (i=0,1,2,3,4); where δ i Represents the standard deviation of the Gaussian kernel of the i-th Gaussian pyramid; i is the pyramid level index; Perform morphological opening operations on each layer of the pyramid to extract multi-scale features and define the feature map of the i-th layer The structural elements The radius r i =3×2 i ;L i Represents the pyramid image of the i-th layer; constructs the low-frequency layer L r With high frequency layer L j Attention weight map A r,j , where j = r + 1, the attention weight map A r,j The formula expression is: In formula (6), ||F r (x,y)||2 is L r The L2 norm feature strength of the layer; For L j The gradient L1 norm of the layer; F r (x,y) is the feature map after the r-th layer morphological opening operation; is the gradient magnitude of the j-th layer image; k is the pyramid level traversal index; The low-frequency continuous features and high-frequency detail features are fused through weighted residual connections to generate an enhanced feature map F. fused , the formula is: In formula (7), A i,i+1 Represents the feature correlation weight between the low-frequency layer i and the high-frequency layer i+1, which is used to dynamically control the contribution of features at different levels in the fusion process; ⊙ represents element-by-element multiplication; ↑ represents the bilinear interpolation upsampling operation, which is used to convert the low-resolution feature map F i Upsample to high resolution; F i+1 Represents the high-frequency feature map of the i+1th layer; Based on the polarized reflection image sequence in step 1, calculate the theoretical value of the refraction angle θ of the scratch area theory , the calculation formula is: In formula (8), n air and n glass are the refractive indices of air and glass, θ inc is the incident angle of the polarized light source in step 1; extract the enhanced feature map F fused The gradient direction θ of the edge point edge , if ||θ theory -θ edge || is greater than the preset angle deviation threshold, it is determined to be an artifact and removed, and the fine screening scratch area map θ is output. refined .

6. The method for detecting scratches on a glass lens surface according to claim 1, wherein: The direction-constrained convolution kernel group extracts the gradient direction histogram of each pixel point based on the fine-screened scratch area map, generates the scratch main direction distribution matrix through Hough transform, and constructs a multi-channel direction-constrained mask based on the refraction angle consistency direction encoding matrix output in step 3. Subsequently, the traditional isotropic convolution kernel is reconstructed into a direction-sensitive kernel structure using a differentiable direction parameterization method. The mask matrix is mapped to the convolution kernel space through the kernel weight direction encoding function to generate a dynamic convolution kernel group within the restricted direction interval, wherein each kernel unit dynamically adjusts the weight distribution within the kernel according to the local direction field. When the direction-constrained convolution kernel group performs the restricted direction convolution operation, a sliding window is used to traverse the fine-screened area map. If the direction encoding value of the window center point exceeds the preset direction interval threshold, the kernel weight zeroing mechanism is triggered to mask the invalid direction response. Finally, the convolution output is texture decomposed by a multi-scale Gabor filter group, and the high-frequency scratch texture is separated from the low-frequency background baseband based on the wavelet packet transform. The multi-scale feature map is spatially projected using the GS orthogonalization algorithm to output the decoupled orthogonal encoding feature vector.

7. The method for detecting scratches on a glass lens surface according to claim 1, wherein: The multi-scale texture decomposition module includes a dual-tree complex wavelet decomposition unit, a direction-sensitive filtering unit, a cross-scale feature fusion unit, a texture orthogonal quantization unit, a non-local similarity aggregation unit, and a spatial orthogonal coding generation unit. The dual-tree complex wavelet decomposition unit is used to decompose the input image into multi-scale high-frequency subbands and low-frequency approximation components through dual-tree complex wavelet transform. The high-frequency subbands cover six directions of ±15°, ±45°, and ±75°, and the low-frequency approximation components retain the global texture contour. The direction-sensitive filtering unit is used to extract the main direction information based on the direction coding matrix DCM and select the high-frequency subband that matches the scratch direction for direction enhancement filtering. The cross-scale feature fusion unit is used to calculate the correlation between adjacent scale subbands using a cross-scale attention mechanism. If the correlation coefficient is higher than a preset threshold, weighted fusion is performed; otherwise, independent scale features are retained. The texture orthogonal quantization unit is used to generate an energy quantization mask by calculating the ratio of the high-frequency subband energy to the low-frequency component. If the high-frequency energy ratio of any area exceeds the preset threshold, it is marked as a scratch candidate area. The non-local similarity aggregation unit is used to perform non-local mean filtering in the energy mask mark area, search for similar texture blocks through a block matching algorithm and perform weighted averaging; the spatial orthogonal coding generation unit is used to perform principal component analysis on the denoised texture map, extract the principal component vector, and splice it with the directional coding matrix DCM to generate an orthogonal coding vector.

8. The method for detecting scratches on a glass lens surface according to claim 1, wherein: The cascade network model includes: a multi-scale feature pyramid input layer, a direction distribution thermal estimation layer, a dynamic FP16 quantization allocation layer, a three-dimensional deformable convolution layer, a residual attention fine inspection layer and a geometric parameter regression layer; the multi-scale feature pyramid input layer is used to generate a multi-scale feature pyramid based on the decoupled feature map input through strided convolution and bilinear interpolation, fuse low-resolution semantic information with high-resolution detail features, and output a pyramid feature tensor; the direction distribution thermal estimation layer is used to perform 3×3 direction-sensitive convolution on the pyramid feature tensor to generate a direction distribution thermal map, and calculate the probability distribution of the scratch direction of each pixel through the Softmax function, and output a thermal confidence matrix; the dynamic FP16 quantization allocation layer is used to trigger FP when the regional confidence of the thermal confidence matrix is greater than or equal to 0.

8. 32 full-precision calculation; when the regional confidence is greater than or equal to 0.5 and less than 0.8, FP16 mixed-precision calculation is enabled; when the regional confidence is less than 0.5, the calculation is skipped, and the quantization mask and the dynamic calculation feature map are output; the three-dimensional deformable convolution layer is used to extract the three-dimensional geometric deformation features of the dynamic calculation feature map through the deformable convolution kernel, and the kernel offset of the deformable convolution kernel is dynamically adjusted by the scratch direction probability to generate a three-dimensional deformation feature tensor; the residual attention fine inspection layer is used to fuse the coarse screening branch and the fine inspection branch features of the three-dimensional deformation feature tensor through the residual connection, and output the fine inspection feature map; the geometric parameter regression layer is used to regress the scratch length, width, depth and curvature parameters of the fine inspection feature map through the fully connected layer and the three-dimensional coordinate mapping, and output a topological structure encoding matrix containing geometric parameters and depth information.

9. The method for detecting scratches on a glass lens surface according to claim 1, wherein: The working principle of step 6 is as follows: The coating process parameter data stream is read in real time through the CAT bus protocol. A multi-physics field coupling model is constructed based on the film thickness and refractive index. The minimum wavelength offset algorithm is used to calculate the compensation light source wavelength adjustment amount Δλ. The calculation formula is: In formula (9), λ0 is the initial compensation light source; n film is the refractive index of the coating layer; n air is the refractive index of air; d is the thickness of the coating layer; Trigger the near-infrared compensation light source in step 1 to switch to λ new =λ0+Δλ; Project the scratch topology output in step 5 to the standard template coordinate system, and extract the offset error matrix E between the actual coordinates and the theoretical coordinates through the Lissajous pattern registration algorithm proj , and generates the incident angle correction Δθ based on the error gradient field inc , the formula is: In formula (10), ||E proj || F Represents the projection error matrix E proj The F norm of the system is used to quantify the overall projection deviation of the system; D base is the reference distance of the standard sample; Dynamically adjust the incident angle of the polarized light source in step 1 to θ′ inc =θ inc +Δθ inc ; At the same time, the frequency domain characteristics of the error matrix are analyzed, and the convolution kernel offset optimization function is used to calculate the deformation vector δ of the three-dimensional deformable convolution kernel in step 4 k , the calculation formula is: In formula (11), Represents the projection error to the convolution kernel weight W k The partial derivative of is used to reflect the sensitivity of the weight to the error; η is the learning rate coefficient; Update the convolution kernel weight W′ k =W k +δ k ; and fuse Δθ through the Bayesian incremental learning algorithm inc and δ k Go to the optical collaborative parameter library to generate a versioned parameter set.

10. A glass lens surface scratch detection system, characterized by: A method for detecting scratches on a glass lens surface as claimed in any one of claims 1 to 9, comprising: a multimodal optical acquisition module, a gradient domain dynamic processing module, a multi-scale feature fusion module, a directional constraint decoupling module, a three-dimensional topological modeling module, and a dynamic closed-loop optimization module; The multimodal optical acquisition module is used to integrate a four-directional polarized light source array and a near-infrared compensation light source, generate a composite light field with orthogonal polarization states through time-sharing triggering, synchronously project the compensation light spot onto the coating surface, and output a four-channel polarized reflection image sequence and a coating interference compensation map; The gradient domain dynamic processing module is used to perform multi-directional gradient convolution on the four-channel polarized reflectance image sequence based on the asymmetric Gabor filter bank, dynamically optimize the segmentation threshold of the improved OTSU algorithm in combination with the noise distribution characteristics of the coating interference compensation image, and output the binary scratch candidate area and the scratch direction encoding matrix; The multi-scale feature fusion module is used to input the binary scratch candidate area into the Gaussian pyramid decomposition network, extract multi-scale morphological features in layers, fuse low-frequency continuous features with high-frequency detail features through a cross-layer attention weighting mechanism, and eliminate artifact interference based on the refraction angle consistency verification algorithm, and output a high-confidence scratch area map after fine screening; The directional constraint decoupling module is used to construct a directional constraint dynamic convolution kernel group based on the directional coding matrix and the finely screened scratch area map, perform multi-scale texture orthogonal decomposition within the restricted directional range, quantize and separate scratch features and base texture noise, and output a decoupled feature map and a spatial orthogonalized code; The three-dimensional topology modeling module is used to input the decoupled feature map into the cascade network model, dynamically allocate FP16 quantization computing resources based on the directional distribution heat map, and fuse the spatial orthogonalization coding to generate a scratch topology structure containing three-dimensional depth information and geometric parameters; The dynamic closed-loop optimization module is used to obtain the coating thickness and refractive index parameters in real time through the CAT bus, perform back-projection error analysis on the scratch topology structure in combination with a standard template, dynamically adjust the compensation light source wavelength of the multimodal optical acquisition module and the convolution kernel offset of the directional constraint decoupling module, and update the light field collaborative parameter library.

Citation Information

Patent Citations

  • Optical lens surface scratch detection method and system

    CN119379620A

  • Glass detection method and system based on artificial intelligence

    CN119804505A

Cited By

  • Circuit board microsection image defect detection method

    CN120635071A

  • Qualitative detection method and device for tiny linear defects on surface of magnetic shoe and medium

    CN120635097A

  • Method for distinguishing cracks and scratches in bonded or annealed LN / LT-SiC wafer

    CN120651855A

  • Unmanned aerial vehicle electric power inspection image intelligent analysis method and system based on deep learning and multi-modal fusion and medium of unmanned aerial vehicle electric power inspection image intelligent analysis method and system

    CN120726041A

  • Table area positioning correction method based on image edge detection

    CN120807564A