Panel flaw detection method and system based on visual cue learning
Patent Information
- Application Number
- CN202610807868.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-06-05
AI Technical Summary
在面板纹理背景呈现高度周期性且瑕疵区域对比度较低的情况下,此类方案所提取的图像特征中背景纹理响应与瑕疵响应相互交织,导致检测模型难以对微弱瑕疵形成有效区分
[0006] This invention transforms the panel image to be detected into the frequency domain and injects a learnable visual cue distribution. It directly differentiates background texture and defect areas at the spectral level, enabling the amplitude cue distribution to suppress energy at frequencies corresponding to periodic background textures, and the phase cue distribution to adjust for phase deviations caused by defect areas. The combined effect significantly weakens background texture and highlights defect areas in the enhanced image after inverse transformation reconstruction. Since the enhancement process occurs at the input front of the pre-trained detection model, all response characteristics of the pre-trained model remain constant, completely avoiding the computational overhead and overfitting tendency caused by fine-tuning the model for production line defect samples. It also retains the fine-grained resolution and generalization performance accumulated by the pre-trained detection model in general visual tasks. After the detection results are generated, this invention constructs a frequency domain cue adjustment based on the current detection results and iteratively corrects the amplitude and phase cue distributions. This allows the visual cue distribution to gradually approach the optimal fit of the spectral characteristics of the panel image during continuous detection, thereby continuously improving the accuracy of panel defect localization and classification without altering the model structure or introducing additional labeled samples.
Smart Images

Figure CN122335873B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and in particular to a method and system for detecting panel defects based on visual cue learning. Background Technology
[0002] Panel defect detection is a technology for identifying and locating surface defects generated during the manufacturing process of display panels. It typically employs an end-to-end detection model based on deep learning, directly inputting the image of the panel to be detected into a pre-trained convolutional neural network. The network extracts image features and outputs the location and category information of the defects. However, when the panel texture background exhibits high periodicity and the defect area has low contrast, the extracted image features from such solutions become intertwined with the background texture response and the defect response, making it difficult for the detection model to effectively distinguish subtle defects. Some solutions attempt to fine-tune the detection model's parameters by collecting defect samples from the production line to adapt to specific panel texture characteristics. However, defect samples in the panel production line occur infrequently and are scattered in form. Obtaining labeled samples covering various defect forms requires significant resources for acquisition and labeling. Furthermore, the fine-tuned model is prone to over-reliance on known panel textures, resulting in a significant decrease in detection stability when facing new panels with slightly different texture characteristics. Therefore, the localization accuracy and classification accuracy of existing panel defect detection solutions in low-contrast defect scenarios need improvement. Summary of the Invention
[0003] This invention provides a method and system for detecting panel defects based on visual cue learning.
[0004] In a first aspect, embodiments of the present invention provide a panel defect detection method based on visual cue learning, comprising: The panel image to be detected is divided into multiple image blocks. A Fourier transform is performed on each image block to obtain the amplitude spectrum representation and phase spectrum representation of the image block. The amplitude spectrum represents the periodic component distribution of the panel texture, and the phase spectrum represents the spatial phase relationship of the panel structure. The pre-constructed amplitude cue distribution and phase cue distribution are invoked. The amplitude cue distribution is applied to the amplitude spectrum representation to selectively suppress the spectral energy corresponding to the background texture, resulting in an enhanced amplitude spectrum representation. At the same time, the phase cue distribution is applied to the phase spectrum representation to adjust the phase distribution of the defect region, resulting in an enhanced phase spectrum representation. Perform inverse Fourier transform on the enhanced amplitude spectrum representation and enhanced phase spectrum representation of each image block to reconstruct the enhanced image block, and then stitch all the enhanced image blocks together according to the division and arrangement order to obtain the complete enhanced image; A pre-trained detection model with fixed input response characteristics of a fully enhanced image is used to perform defect localization and category prediction on the fully enhanced image and output the current detection result, which includes defect location and defect category labels. Based on the current detection results, a frequency domain cue adjustment amount is constructed. The amplitude cue distribution and phase cue distribution are corrected by the frequency domain cue adjustment amount to obtain the corrected amplitude cue distribution and the corrected phase cue distribution. Based on the corrected amplitude cue distribution and the corrected phase cue distribution, the process of calling the pre-constructed amplitude cue distribution and phase cue distribution to output the current detection result is repeated until the cue adjustment termination condition is met. The defect location mark and defect category mark of the last output are used as the final defect detection information.
[0005] Secondly, embodiments of the present invention provide a computer system including a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement the steps in the above method.
[0006] This invention transforms the panel image to be detected into the frequency domain and injects a learnable visual cue distribution. It directly differentiates background texture and defect areas at the spectral level, enabling the amplitude cue distribution to suppress energy at frequencies corresponding to periodic background textures, and the phase cue distribution to adjust for phase deviations caused by defect areas. The combined effect significantly weakens background texture and highlights defect areas in the enhanced image after inverse transformation reconstruction. Since the enhancement process occurs at the input front of the pre-trained detection model, all response characteristics of the pre-trained model remain constant, completely avoiding the computational overhead and overfitting tendency caused by fine-tuning the model for production line defect samples. It also retains the fine-grained resolution and generalization performance accumulated by the pre-trained detection model in general visual tasks. After the detection results are generated, this invention constructs a frequency domain cue adjustment based on the current detection results and iteratively corrects the amplitude and phase cue distributions. This allows the visual cue distribution to gradually approach the optimal fit of the spectral characteristics of the panel image during continuous detection, thereby continuously improving the accuracy of panel defect localization and classification without altering the model structure or introducing additional labeled samples. Attached Figure Description
[0007] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present invention and, together with the specification, serve to explain the technical solutions of the present invention.
[0008] Figure 1 This is a schematic diagram illustrating the principle of a panel defect detection method based on visual cue learning, provided in an embodiment of the present invention.
[0009] Figure 2 This is a schematic diagram illustrating the implementation process of a panel defect detection method based on visual cue learning, provided in an embodiment of the present invention.
[0010] Figure 3 This is a schematic diagram of the hardware entity of a computer system provided in an embodiment of the present invention. Detailed Implementation
[0011] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0012] This invention provides a panel defect detection method based on visual cue learning, which can be executed by a computer system processor. The computer system can refer to a device with data processing capabilities, such as a server, laptop, tablet, or desktop computer.
[0013] Please combine Figure 1Referring to the present invention, the panel defect detection method based on visual cue learning can be summarized into several key steps: frequency domain enhancement, spatial domain reconstruction, model detection, and adaptive iteration of cues. Specifically, the panel image to be detected is first divided into image blocks and transformed to the frequency domain, simultaneously acquiring the amplitude spectrum recording the distribution of periodic components and the phase spectrum recording the spatial phase relationship. Utilizing the amplitude and phase cue distributions obtained through statistical learning from a large number of defect-free samples, the system can selectively suppress the spectral energy of the panel background texture in the frequency domain and adjust the phase structure of defect-prone areas through perturbation. The enhanced spectral information is reconstructed back to the spatial domain through inverse transformation, and then, after dynamic range stretching and seam transition fusion, it is stitched together into a complete enhanced image where the background is weakened and defects are highlighted. This enhanced image is then fed into a pre-trained detection model with fixed parameters. The model uses its hierarchical deep convolutional processing flow to simultaneously complete the activation and localization of suspected defect areas, accurate bounding box regression, and multi-class defect probability mapping, thereby outputting the defect location and category label for the current round. The detection results are not the final output, but rather serve as feedback signals to drive an automatic correction of the frequency domain cues: by comparing the spectrum of the detected defective region with the ideal defect-free spectrum model, the system generates amplitude and phase cues errors, respectively, and uses these errors to perform directional superposition correction on the original amplitude and phase cues distribution. The updated cues distribution is then applied back to the original spectrum, undergoing reconstruction, stitching, and model inference again to generate a new round of detection results, thus forming an iterative closed loop. The iteration terminates when the change in the intersection-union ratio of all defect bounding boxes between adjacent rounds is lower than a preset convergence limit, and the output defect location and category labels at this point are the final detection information. The method of this invention deeply integrates the frequency domain manipulation capability of visual cue learning with the deep semantic extraction capability of the detection model, and through an iterative self-optimization mechanism for the cues distribution, continuously transforms the detection feedback into a refined correction of frequency domain suppression and perturbation strategies, thereby achieving stable capture and fine segmentation of minute defects under complex texture interference, significantly improving the accuracy and robustness of panel defect detection.
[0014] The following is a detailed introduction for reference. Figure 2 This is a schematic diagram illustrating the implementation process of the panel defect detection method based on visual cue learning provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the method includes the following steps S100~S500: Step S100: Divide the panel image to be detected into multiple image blocks, perform Fourier transform on each image block to obtain the amplitude spectrum representation and phase spectrum representation of the image block. The amplitude spectrum represents the periodic component distribution of the panel texture, and the phase spectrum represents the spatial phase relationship of the panel structure.
[0015] The panel image to be inspected is a single-channel image obtained by converting a grayscale or color image of the panel product surface, acquired under standard lighting conditions using an industrial linear or area scan camera, into grayscale. The image contains background texture information and potential defect information. An image block is a rectangular sub-image unit created by cutting the entire panel image according to a preset grid division rule. Each image block serves as the basic processing unit for subsequent frequency domain transformation. The division operation makes the texture distribution within each image block tend to be locally stable, which is beneficial for the accurate extraction of periodic components during frequency domain analysis. The Fourier transform, for example, is a two-dimensional discrete Fourier transform. This transform maps the image block from the spatial domain to the frequency domain, decomposing it into a linear combination of sinusoidal elementary signals with different frequencies, directions, amplitudes, and initial phases. This converts the spatial arrangement information of pixel grayscale values into complex representations of each frequency point in the frequency domain. The amplitude spectrum is a real matrix formed by taking the modulus of each element in the complex matrix obtained from the Fourier transform. The amplitude value at each frequency point represents the proportion of energy of that frequency component in the image patch. The larger the amplitude value, the more significant the periodic texture corresponding to that frequency is in the image. Panel textures are usually represented by regular stripes or grid structures that repeat along a specific direction, forming concentrated energy response peaks in the corresponding direction on the amplitude spectrum. Therefore, the amplitude spectrum represents the distribution of the periodic components of the panel texture. The phase spectrum is a real matrix formed by taking the principal argument value of each element in the complex matrix obtained from the Fourier transform. The phase value at each frequency point represents the initial offset of that frequency component during spatial synthesis, determining the specific spatial position of the sine wave of that frequency. The spatial positions of geometric features such as boundaries, edges, and contours of different regions in the panel structure are jointly determined by the phase relationship between different frequency components. Therefore, the phase spectrum represents the spatial phase relationship of the panel structure.
[0016] When dividing the image into blocks, the panel image to be detected is divided into grids according to a preset standard block size. The standard block size is determined based on the periodic scale of the panel texture and the detection resolution requirements. For areas where the image edges are insufficient to form a complete standard block size, pixels are supplemented by symmetrical extension outward along the image boundary. During extension, the pixel values inside the boundary are mirrored and copied to the extension area with the boundary as the axis. This ensures that the image blocks formed after extension maintain texture continuity similar to the internal areas at the boundaries, avoiding the introduction of artificial high-frequency components due to differences in edge fill values that could interfere with the subsequent Fourier transform results. After the panel image to be detected is divided, a two-dimensional discrete Fourier transform is performed sequentially on each image block.
[0017] Specifically, for the grayscale values of each pixel within the image patch, a one-dimensional discrete Fourier transform is first performed row by row along the row direction, converting the spatial index of the row direction into a horizontal frequency index, resulting in an intermediate complex matrix. Then, a one-dimensional discrete Fourier transform is performed column by column along the column direction of the intermediate complex matrix, converting the spatial index of the column direction into a vertical frequency index, ultimately obtaining the complete complex spectral representation of the image patch. The computational mechanism of the one-dimensional discrete Fourier transform is as follows: for an input sequence of length L, at each frequency index k, the nth element in the input sequence is multiplied and accumulated by a rotation factor with an angle of -2π×k×n / L. All components from n to L-1 are accumulated to obtain the complex spectral value at that frequency index. After completing the two-dimensional discrete Fourier transform, a quadrant centering shift operation is performed on the complex spectrum, diagonally swapping the four quadrants so that the zero-frequency component is located at the center of the spectral plane, the low-frequency components are distributed around the center, and the high-frequency components are located at the edges of the spectral plane, facilitating the location of the effective areas of the subsequent amplitude cue distribution and phase cue distribution. The process of extracting the amplitude spectrum representation from the quadrant-centered complex spectrum is as follows: For each frequency point in the complex spectrum, the magnitude of the stored complex value is taken as the amplitude value at that frequency point. The magnitude is calculated by multiplying the real part of the complex value by itself, then multiplying the imaginary part by itself, and taking the square root, resulting in an amplitude matrix of the same size as the frequency domain plane. The process of extracting the phase spectrum representation from the quadrant-centered complex spectrum is as follows: For each frequency point in the complex spectrum, the argument of the stored complex value is taken as the phase value at that frequency point. The argument is calculated by dividing the imaginary part of the complex value by the arctangent of the ratio corresponding to the real part. The quadrant is determined based on the signs of the real and imaginary parts, and the phase value is reduced to the principal value interval from negative π to positive π, resulting in a phase matrix of the same size as the frequency domain plane. After the above processing, each image block obtains its own amplitude spectrum representation and phase spectrum representation. The frequency domain information carried by these two representations is complementary, together forming a complete frequency domain characterization of the image block.
[0018] Step S200: Call the pre-constructed amplitude cue distribution and phase cue distribution, apply the amplitude cue distribution to the amplitude spectrum representation to selectively suppress the spectral energy corresponding to the background texture, and obtain an enhanced amplitude spectrum representation. At the same time, apply the phase cue distribution to the phase spectrum representation to adjust the phase distribution of the defect area, and obtain an enhanced phase spectrum representation.
[0019] The amplitude cue distribution is a two-dimensional weighted matrix with the same frequency-domain planar size as the amplitude spectrum representation of the image patch. Each element in this matrix records the amplitude adjustment indication at the corresponding frequency point. The magnitude of the indication determines the degree to which the spectral energy at that frequency point is preserved or suppressed. Its construction process is based on frequency-domain statistical learning of a set of defect-free panel images. The phase cue distribution is a two-dimensional offset matrix with the same frequency-domain planar size as the phase spectrum representation of the image patch. Each element in this matrix records the phase offset indication at the corresponding frequency point. The magnitude of the offset determines the magnitude of the phase value adjustment at that frequency point. Its construction process is based on statistical analysis of the phase stability of a set of defect-free panel images. The spectral energy corresponding to the background texture is the concentrated energy response formed by the regularly repeating texture pattern on the panel surface in the amplitude spectrum. This energy is concentrated at specific frequency points, causing the panel background texture to exhibit a significant periodic appearance in the spatial domain. In defect detection tasks, this background texture energy can mask the weak frequency-domain perturbations caused by defects. The phase distribution of defective regions represents the non-periodic structural features of the panel surface in the phase spectrum. Due to their irregular shapes and abrupt boundary changes, defective regions exhibit phase response patterns different from the background texture in the phase spectrum. The enhanced amplitude spectrum representation is an optimized amplitude matrix obtained by applying amplitude cue distribution to the original amplitude spectrum. In this matrix, the spectral energy corresponding to the background texture is selectively reduced, while the spectral energy of non-background texture frequency bands is preserved or even relatively highlighted. The enhanced phase spectrum representation is an optimized phase matrix obtained by applying phase cue distribution to the original phase spectrum. In this matrix, the phase of defective regions is targeted and adjusted, making the structural features of the defects more prominent during subsequent inverse transform reconstruction.
[0020] In one embodiment, step S200 may specifically include the following steps S210 to S260: Step S210: Obtain multiple standard images of defect-free panels, perform Fourier transform on each standard image of defect-free panels, and obtain the standard amplitude spectrum set and the standard phase spectrum set.
[0021] The standard images of defect-free panels are sample images of panel products that have been visually inspected or confirmed by higher-precision testing equipment to be free of any defects. These images contain only the normal background texture and structural information of the panel, and do not contain scratches, dents, dirt, cracks, or other defects. The standard amplitude spectrum set is an amplitude spectrum sample database compiled by sequentially performing the same image segmentation, Fourier transform, and amplitude extraction operations as in step S100 on each of the standard images of defect-free panels, and then aggregated in the frequency domain plane. Each frequency point corresponds to the amplitude value records of multiple defect-free images. The standard phase spectrum set is a phase spectrum sample database compiled by sequentially performing the same image segmentation, Fourier transform, and phase extraction operations as in step S100 on each of the standard images of defect-free panels, and then aggregated in the frequency domain plane. Each frequency point corresponds to the phase value records of multiple defect-free images.
[0022] The methods for obtaining multiple defect-free panel standard images are as follows: retrieving confirmed defect-free product images from the historical inspection database of the panel production line, or specifically collecting a batch of rigorously inspected and qualified panel samples at the production site and taking images. The number of samples must be sufficient to cover the possible texture variation range and illumination fluctuation range of the panel model. The process of performing Fourier transform on each defect-free panel standard image is consistent with the processing flow in step S100: first, the image is divided into image blocks of the same standard block size; then, a two-dimensional discrete Fourier transform is performed on each image block, followed by quadrant centering shift, and the amplitude spectrum and phase spectrum are extracted. Image blocks at the same grid position in all defect-free panel standard images constitute an image block group. All amplitude spectra corresponding to this image block group are collected into a standard amplitude spectrum set for that image block group position. The amplitude value sequence of each image in the group is stored at each frequency point position in the standard amplitude spectrum set. Similarly, all phase spectra corresponding to this image block group are collected into a standard phase spectrum set for that image block group position. The phase value sequence of each image in the group is stored at each frequency point position in the standard phase spectrum set. Each image patch group has its own independent set of standard amplitude spectrum and standard phase spectrum. The subsequent construction of amplitude cue distribution and phase cue distribution is performed separately for each image patch group to adapt to the subtle differences in texture in different areas of the panel.
[0023] Step S220: Perform frequency-point aggregation processing on the standard amplitude spectrum set to obtain the standard amplitude spectrum aggregation result. Identify the frequency points in the standard amplitude spectrum aggregation result whose energy concentration exceeds the preset limit, and determine the identified frequency points as the main frequency point set of the background texture.
[0024] Frequency-by-frequency aggregation processing calculates the central tendency representative value for each frequency point in the standard amplitude spectrum set using a statistical aggregation function. This results in a single numerical value for each frequency point that comprehensively reflects the typical amplitude level of multiple defect-free images at that frequency component. The standard amplitude spectrum aggregation result is a real matrix with the same frequency-domain plane size as the image patch amplitude spectrum. Each element in the matrix stores the amplitude representative value calculated for the corresponding frequency point. The energy concentration degree is the prominence of the amplitude representative value at a certain frequency point in the standard amplitude spectrum aggregation result relative to the overall distribution level of the amplitude representative values in the entire frequency-domain plane. A higher amplitude representative value indicates that the texture component corresponding to that frequency point is prevalent and strongly present in multiple defect-free images, constituting the intrinsic frequency component of the panel background texture. The preset limit is a discrimination threshold determined based on the statistical distribution of the amplitude representative values in the frequency-domain plane, used to distinguish the intrinsic frequency component of the background texture from general spectral floor noise. The set of main frequency points of the background texture is a coordinate index set composed of all frequency points whose energy concentration exceeds a preset limit. The positions of these frequency points identify the characteristic response positions of the panel background texture in the frequency domain, and serve as the core suppression target area in the subsequent amplitude suppression distribution construction.
[0025] The specific method for frequency-point aggregation of the standard amplitude spectrum set is as follows: Take the amplitude value sequence at each frequency point and calculate the median of the sequence as the representative amplitude value at that frequency point. The median is chosen as the aggregation function because it is insensitive to a small number of abnormal amplitude deviations and can robustly reflect the typical energy level of that frequency point under defect-free conditions. This avoids the influence of extreme amplitude values introduced by slight lighting differences or subtle texture variations in individual images on the aggregation results. After obtaining the standard amplitude spectrum aggregation results, the process of identifying frequency points where the energy concentration exceeds a preset limit is as follows: First, arrange the representative amplitude values of all frequency points in the standard amplitude spectrum aggregation results in descending order, and use the representative amplitude value located at a preset percentile position in this arrangement as the preset limit. The selection of the preset percentile should ensure that significant background texture frequencies are fully included while avoiding the inclusion of too much frequency domain noise. After determining the preset limit, each frequency point in the standard amplitude spectrum aggregation result is traversed. If the amplitude representative value of the frequency point is greater than the preset limit, the coordinate index of the frequency point is recorded in the background texture main frequency point set; if the amplitude representative value of the frequency point is less than or equal to the preset limit, the frequency point is excluded from the background texture main frequency point set. The distribution of the background texture main frequency point set on the frequency domain plane usually presents a symmetrical point or strip clustering pattern, and its distribution shape directly corresponds to the directionality and period length of the spatial domain panel texture.
[0026] Step S230: Construct an amplitude suppression distribution based on the set of dominant frequency points of the background texture. The amplitude suppression distribution has a response amount less than the reference flux at each frequency point within the set of dominant frequency points of the background texture, and has a response amount of the reference flux at frequency points outside the set of dominant frequency points of the background texture. The constructed amplitude suppression distribution is used as the amplitude cue distribution.
[0027] The amplitude suppression distribution is a two-dimensional response matrix with the same frequency-domain plane size as the image patch amplitude spectrum. Each element in the matrix stores a response value, which represents the multiplicative coefficient for spectral energy adjustment at the corresponding frequency point in the amplitude spectrum. A response value equal to the baseline flux indicates that the spectral energy at that frequency point passes through with its original amplitude value without any suppression; a response value less than the baseline flux indicates that the spectral energy at that frequency point is compressed and weakened, with smaller values indicating stronger weakening. The baseline flux is a preset reference response value, typically the response level corresponding to the full-pass state, indicating complete preservation of spectral energy. The amplitude cue distribution applies a response value lower than the baseline flux at each frequency point within the main frequency point set of the background texture, causing the spectral energy of the panel background texture at these frequency points to be attenuated after the effect, thereby weakening the representation intensity of the background texture in the spatial domain reconstructed image; at frequency points outside the main frequency point set of the background texture, the response value of the baseline flux is maintained, allowing frequency domain components that do not belong to the background texture to be preserved as is, and non-periodic frequency domain responses caused by defects to be unintentionally suppressed.
[0028] In one embodiment, step S230 may specifically include the following steps S231 to S236: Step S231: Map each frequency point in the main frequency point set of the background texture to the corresponding suppression control unit in the amplitude suppression distribution, and assign an initial suppression degree to each suppression control unit.
[0029] The suppression control unit is an element located at the coordinates of the dominant frequency point of the background texture in the amplitude suppression distribution matrix. Each suppression control unit is responsible for suppressing the spectral energy at that frequency point. The initial suppression level is the initial response value assigned to the suppression control unit before any fine-tuning. This value is less than the reference flux, indicating that the spectral energy at that frequency point is initially attenuated to a certain extent. The mapping relationship between the suppression control unit and the dominant frequency point of the background texture is a one-to-one spatial location mapping, that is, the frequency domain coordinates of the dominant frequency point of the background texture directly correspond to the element position of the same row and column index in the amplitude suppression distribution matrix. The process of assigning the initial suppression level is as follows: a uniform initial response value is set, which is used as the starting point for adjustment of all suppression control units; the magnitude of the initial response value should be such that the background texture is effectively suppressed without causing excessive loss of spectral information, and it can be optionally set to half of the reference flux.
[0030] Step S232: Perform a main frequency energy attenuation trend analysis on the standard amplitude spectrum aggregation results. Based on the attenuation trend of energy at each background texture main frequency point with frequency change, adjust the suppression degree of the corresponding suppression control unit so that the suppression degree is correlated with the attenuation trend in the same direction.
[0031] The analysis of the dominant frequency energy attenuation trend examines the spatial variation of the amplitude representation value of the standard amplitude spectrum aggregation result in the surrounding frequency domain for each dominant frequency point of the background texture. In particular, it analyzes the deceleration rate and decreasing shape of the amplitude representation value as it extends away from the dominant frequency point, thereby revealing the steepness or gentleness of the energy concentration of the background texture component in the frequency space. The attenuation trend characterizes the dispersion range of the spectral energy of the dominant frequency point of the background texture in the frequency space. If the amplitude representation value around the dominant frequency point decreases rapidly with increasing distance, it indicates that the texture component has high energy concentration and high frequency purity; if the amplitude representation value around the dominant frequency point decreases slowly with increasing distance, it indicates that the energy distribution of the texture component is relatively dispersed and has a wide frequency range. The degree of suppression is correlated with the attenuation trend. For the main frequency point with rapid energy attenuation and high frequency concentration, a relatively light degree of suppression is applied because its energy distribution range is narrow, and an excessive suppression range will affect the useful frequency components in its neighborhood. For the main frequency point with slow energy attenuation and high frequency dispersion, a relatively heavy degree of suppression is applied because its energy radiation range is wide, and stronger suppression is required to effectively reduce the frequency domain contribution of background texture.
[0032] During the analysis of the main frequency energy decay trend, for each frequency point in the set of main frequency points of the background texture, its neighboring frequency points are taken radially outward along the frequency domain plane with that frequency point as the center. The radius of the neighboring range is selected to be adapted to the frequency domain resolution of the panel texture. The average value of the amplitude representative value at each distance layer from the center outward is calculated radially, forming a decay curve with distance as the independent variable and the average amplitude representative value as the dependent variable. A monotonically decreasing function is fitted to this decay curve, and the absolute value of the negative slope at the center point is calculated. The larger the absolute value, the faster the decay; the smaller the absolute value, the smoother the decay. When adjusting the suppression level of the corresponding suppression control unit, the initial suppression level is weighted and corrected according to the magnitude of the absolute value of the negative slope: for main frequency points with a large absolute value of the negative slope, the initial suppression level is adjusted back towards the reference flux by a certain amount, reducing its suppression level; for main frequency points with a small absolute value of the negative slope, the initial suppression level is further adjusted downward away from the reference flux by a certain amount, increasing its suppression level. This adjustment ensures that the amplitude suppression at each major frequency point is consistent with the frequency domain spatial distribution characteristics of the background texture, thus avoiding the problems of insufficient suppression of narrowband textures and excessive suppression of wideband textures.
[0033] Step S233: Perform neighborhood smoothing on the adjusted suppression control unit, and harmonize the suppression degree of each suppression control unit with the suppression degree of the suppression control unit at its adjacent frequency point so that the suppression degree transitions continuously in the frequency domain.
[0034] Neighborhood smoothing involves using spatial filtering to weighted average the suppression levels of each suppression control unit within the set of dominant background texture frequencies in the amplitude suppression distribution matrix. This weighted average of the suppression levels of each control unit and its spatial neighbors creates a smoother, more continuous change in suppression level across the frequency domain, eliminating abrupt changes in suppression levels between isolated control units and their surrounding frequencies. This continuous transition in suppression levels in the frequency domain ensures a smooth variation in the spectral energy modulation generated when the amplitude cue distribution acts on the amplitude spectrum. This prevents the introduction of artificial ringing effects or artifacts caused by abrupt changes in spectral energy during subsequent inverse Fourier transform reconstruction of the spatial domain image.
[0035] For neighborhood smoothing, a two-dimensional smoothing window is defined, covering a frequency-domain neighborhood centered on the current suppression control unit. The current suppression level values of all suppression control units within this window that belong to the set of dominant background texture frequencies are multiplied by their corresponding smoothing weights and then summed. The smoothing weights are determined based on the spatial distance from the frequency point within the window to the center frequency point; the closer the distance, the greater the weight, and the farther the distance, the smaller the weight. The sum of the weights is normalized to keep the suppression level values within a reasonable range after smoothing. The sum is used as the new suppression level for that center frequency point after smoothing. This operation is performed sequentially for each suppression control unit to obtain the suppression level distribution after neighborhood smoothing. For suppression control units located at the edge of the set of dominant background texture frequencies, some positions within their smoothing window belong to non-dominant frequency regions. These positions do not participate in the weighted average smoothing; only the effective dominant frequencies within the window are used in the calculation, ensuring that the smoothing operation only occurs within the set of dominant frequencies.
[0036] Step S234: Set up an all-pass control unit at the non-dominant frequency position of the amplitude suppression distribution, and assign a reference flux value to the all-pass control unit so that the amplitude at the non-dominant frequency position can be preserved.
[0037] Non-dominant frequency positions are all coordinate positions on the frequency domain plane that are outside the set of dominant frequency points of the background texture. The energy concentration of these frequency points in the standard amplitude spectrum aggregation result does not exceed the preset limit, and they do not belong to the characteristic frequencies of the panel background texture. Therefore, the corresponding spectral energy should be completely preserved. The all-pass control unit is the element located at the non-dominant frequency position in the amplitude suppression distribution matrix. Its function is to ensure that the spectral energy at this frequency point is not attenuated after the amplitude cue distribution is applied. The reference flux value is a response constant representing the all-pass state. Assigning this constant to the all-pass control unit means that the original amplitude value at this frequency point is transmitted intact to the enhanced amplitude spectrum representation after the application.
[0038] After calculating the suppression level and performing neighborhood smoothing on the suppression control units within the set of dominant background texture frequencies, the entire amplitude suppression distribution matrix is traversed. For positions not belonging to the set of dominant background texture frequencies, the element type at that position is marked as an all-pass control unit, and its response value is directly assigned as the baseline flux value. The baseline flux value serves as the reference baseline for the response values throughout the amplitude suppression distribution matrix; the response values of suppression control units are all less than this baseline, and the response values of all-pass control units are all equal to this baseline.
[0039] Step S235: Combine all suppression control units and all-pass control units into an overall amplitude suppression distribution. The response of the overall amplitude suppression distribution changes continuously in the frequency domain and the central area corresponds to the set of main frequency points of the background texture. This overall amplitude suppression distribution is determined as the amplitude prompt distribution.
[0040] The overall amplitude suppression distribution is a complete two-dimensional response matrix assembled from the suppression control unit (which underwent suppression degree calculation and smoothing in steps S232 and S233) and the full-pass control unit (set as the reference flux value in step S234), according to their respective frequency domain coordinate positions. The assembled overall amplitude suppression distribution exhibits the following characteristics in the frequency domain plane: in local regions centered on each background texture dominant frequency point, the response quantity continuously transitions from a lower value at the center to the reference flux value towards the periphery, forming a smooth funnel-shaped concave distribution of the response quantity in the frequency domain space; in frequency domain regions far from the set of background texture dominant frequency points, the response quantity uniformly maintains the reference flux value. The central region corresponds to the set of background texture dominant frequency points, where the geometric center of each response quantity concave point in the overall amplitude suppression distribution coincides with the coordinate position of the background texture dominant frequency point. The spatial range and steepness of the concave points are determined by the combined effect of dominant frequency energy attenuation trend analysis and neighborhood smoothing.
[0041] Step S236: Perform frequency domain range constraint processing on the amplitude prompt distribution, shrink the frequency domain span of the suppression transition band while maintaining the degree of suppression of the main frequency point, so that the suppression transition region of the amplitude prompt distribution is concentrated around the set of main frequency points of the background texture.
[0042] Frequency domain range constraint processing remaps the response values in the transition region between the suppression control unit and the full-pass control unit in the amplitude cue distribution. By tightening the frequency domain width of the transition band, the suppression effect is more concentrated on the immediate region of the dominant frequency point set of the background texture, reducing the indirect impact of the suppression operation on the frequency domain components of non-background textures. The suppression transition band is the frequency domain space range through which the response value in the amplitude cue distribution gradually climbs from the lowest value at the dominant frequency point to the reference flux value. The response value within this range is below the reference flux but not the lowest value. Shrinking the frequency domain span of the suppression transition band compresses the distance required for the transition band to climb outward from the center to the reference flux while keeping the response value at the center of each suppression control unit constant, allowing the suppression effect to switch from the strongest suppression to complete cancellation of suppression within a narrower frequency domain range.
[0043] One approach to performing frequency domain range constraint processing involves establishing a response profile radially around each suppression control unit in the amplitude cue distribution, with that unit as the center. This profile records the response values at each frequency point from the center to the outer edge of the transition band. A monotonic nonlinear transformation is applied to this profile. The transformation function maintains its original value in the low response region near the center, rapidly increases the response value to approach the reference flux in the intermediate transition region, and maintains the reference flux unchanged far beyond the transition region. This nonlinear transformation can be implemented using a function with a steep rise characteristic, such as a centrally symmetric S-shaped function, whose control parameters are set to shorten the radial length of the rising segment of the transition region to a preset proportion of the original transition region length. This radial profile transformation operation is repeated for each suppression control unit in the amplitude cue distribution to obtain the frequency domain range-constrained amplitude cue distribution. After processing, the suppression transition region of the amplitude cue distribution is compressed into a narrow annular region tightly surrounding each background texture main frequency point, significantly enhancing the frequency selectivity of the suppression operation.
[0044] Step S240: Perform phase deviation statistics on the standard phase spectrum set, determine the phase fluctuation range of each frequency point position between the defect-free panel standard images, determine the frequency points whose phase fluctuation range is less than the preset stability limit as the phase stable frequency point set, generate a phase perturbation distribution based on the phase stable frequency point set, the phase perturbation distribution carries a preset offset at the frequency points outside the phase stable frequency point set, and use the generated phase perturbation distribution as the phase prompt distribution.
[0045] Phase deviation statistics calculate the dispersion or range of phase values among multiple defect-free images stored at each frequency point in a standard phase spectrum set. This measures the consistency of the phase at that frequency point across multiple normal samples. Phase fluctuation range is a quantitative description of the range of phase value variation at a frequency point. It is typically expressed as the difference between the maximum and minimum values of the phase value sequence at that frequency point, or as a certain quantile interval of the phase value sequence at that frequency point. A larger fluctuation range indicates that the phase at that frequency point is more unstable across different normal samples. A preset stability limit is a fluctuation range threshold used to distinguish between phase-stable and phase-unstable frequency points. Frequency points with fluctuation ranges smaller than this threshold are considered to have highly consistent phase performance in the normal panel, constituting a reliable carrier of structural information. Frequency points with fluctuation ranges greater than or equal to this threshold are considered to have drastic phase changes across different normal samples, lacking the ability to stably represent structural information, but precisely representing the operational region where defects may introduce phase anomalies. The phase-stable frequency set is composed of the coordinates of all frequency points with fluctuation ranges smaller than a preset stability limit. The phase of these frequency points remains stable within the normal panel, reflecting the spatial phase relationship of the panel's inherent geometry. The phase perturbation distribution is a two-dimensional offset matrix with the same frequency-domain planar size as the image patch phase spectrum. Each element in the matrix records the additional offset applied to the original phase value at the corresponding frequency point. The magnitude and sign of the offset determine the direction and amplitude of the phase adjustment. The phase perturbation distribution carries a preset offset at frequency points outside the phase-stable frequency set, indicating that phase adjustment is only performed on frequency regions with unstable phases within the normal panel, while maintaining the phase integrity of stable regions.
[0046] In one embodiment, step S240 may specifically include the following steps S241 to S246: Step S241: Obtain the phase values of each defect-free panel standard image in the standard phase spectrum set at the same frequency point, and obtain the phase value sequence of each corresponding frequency point.
[0047] The standard phase spectrum set already stores all phase spectrum data according to three dimensions: image patch group location, defect-free image number, and frequency point coordinates. The process of obtaining the phase values of each defect-free panel standard image at the same frequency point is as follows: for a specified image patch group location, traverse all frequency point coordinates in the frequency domain plane, and extract the stored phase values of all images in that group at each frequency point coordinate, forming a phase value sequence with a length equal to the number of defect-free images in that group. During extraction, attention must be paid to the cyclic nature of the phase values; that is, the phase values are distributed within the main value range of -π to +π. Phase values close to +π and close to -π differ significantly numerically but have a very small physical phase difference. Therefore, a cyclic reading method is needed to keep the original records of the phase values unchanged so that appropriate cyclic statistical algorithms can be used for subsequent statistical processing.
[0048] Step S242: Perform discreteness statistics on the phase value sequence corresponding to each frequency point to obtain the phase discreteness quantity. Mark the frequency points with the phase discreteness quantity less than the preset stability threshold as phase stable frequency points, and collect all phase stable frequency points to form a phase stable frequency point set.
[0049] Because phase values exhibit cyclic wrapping characteristics, ordinary statistical methods for dispersion based on linear scales cannot be directly applied to phase data. Therefore, cyclic variance or cyclic standard deviation is used as a measure of phase dispersion. Specifically, each phase value in the phase value sequence is first mapped to a unit vector on a unit circle. The horizontal component of this unit vector is the cosine of the phase value, and the vertical component is the sine of the phase value. The average vector is then calculated for all unit vectors. The horizontal component of the average vector is the mean of the horizontal components of all vectors, and the vertical component is the mean of the vertical components of all vectors. The magnitude of the average vector is called the length of the cyclic mean composite vector. This length ranges from zero to one. The closer the length is to one, the higher the concentration of the phase value sequence; the closer the length is to zero, the higher the dispersion of the phase value sequence. The measure of phase dispersion is defined as the complement of the length of the cyclic mean composite vector, i.e., one minus the length of the composite vector. The smaller the value, the more concentrated the phase; the larger the value, the more dispersed the phase. The phase dispersion measure is compared with a preset stability threshold, which is a pre-selected positive decimal determined by referring to the statistical quantile of the complement of the resultant vector length of the most phase-stable frequency point in the normal panel. If the phase dispersion measure is less than the preset stability threshold, it indicates that the phase of that frequency point is highly consistent among multiple defect-free images, and the frequency point is marked as a phase-stable frequency point. If the phase dispersion measure is greater than or equal to the preset stability threshold, it indicates that the phase of that frequency point fluctuates greatly and is not included in the set of phase-stable frequency points. All frequency points in the frequency domain plane are traversed, and the coordinates of all phase-stable frequency points are collected to form a set of phase-stable frequency points.
[0050] Step S243: Determine the spatial connected domains of the set of phase-stable frequency points in the frequency domain plane, perform extrapolation and expansion processing on each spatial connected domain to obtain the phase protection zone boundary. The frequency points inside the phase protection zone boundary maintain the original phase relationship, while the frequency points outside the boundary undergo phase shift processing.
[0051] The frequency points in the set of phase-stable frequency points are not isolated and scattered on the frequency domain plane, but rather form several connected regions due to the continuity of the spatial frequency of the panel structure. These connected regions are called spatially connected domains. The process of determining spatially connected domains adopts a binary image connected component labeling algorithm: the frequency domain plane is regarded as a binary mesh graph, phase-stable frequency points are labeled as foreground, and unstable frequency points are labeled as background. The mesh is traversed using the four-connectivity or eight-connectivity criterion, and adjacent connected foreground pixels are classified into the same connected domain and assigned a unique domain label. After obtaining each spatially connected domain, an extrapolation dilation process is performed on each connected domain. The extrapolation dilation process expands the original geometric boundary of the connected domain outward along the normal direction by a preset number of layers. The number of expansion layers is related to the total size of the frequency domain plane and the texture period. The expanded region contains the original connected region and a ring-shaped or strip-shaped extended region layer around it. The specific implementation of the extrapolation expansion process is as follows: starting from the boundary frequency of the original connected domain, the process explores outward along the eight principal directions of the frequency domain coordinate axis. All frequency points located outside the original connected domain and within a specified distance from the boundary frequency point are included in the expansion region. The distance is set so that the expansion region contains a transition space for adjacent frequency phases in a defect-free state. The boundary of the phase protection zone is defined by the outer boundary of the connected domain after the extrapolation expansion process. All frequency points inside the boundary maintain the original phase relationship, that is, the positions corresponding to these frequency points are filled with zero offset in the phase perturbation distribution; all frequency points outside the boundary undergo phase offset processing, that is, the positions corresponding to these frequency points are filled with a non-zero preset offset in the phase perturbation distribution.
[0052] Step S244: In the phase disturbance distribution, a preset offset is filled in for each frequency point outside the phase protection zone boundary. The amplitude of the preset offset is monotonically increasing with the shortest distance from the frequency point to the phase protection zone boundary.
[0053] The preset offset is a phase offset value assigned to frequency points belonging to non-phase protection zones in the phase perturbation distribution matrix. This offset value is directly superimposed on the original phase value of the corresponding frequency point when the phase hint distribution is invoked. To ensure a smooth spatial transition of the phase offset and avoid artificial structural artifacts caused by phase discontinuities introduced at the boundary of the phase protection zone due to abrupt shifts in offset, the magnitude of the preset offset is designed to monotonically increase with the shortest distance from the frequency point to the boundary of the phase protection zone. That is, the farther the frequency point is from the boundary, the greater the magnitude of the phase adjustment, and the offset of the frequency point immediately adjacent to the boundary approaches zero, achieving a smooth transition of the offset from zero offset inside the protection zone to full offset outside. The shortest distance is measured using Euclidean distance in the frequency domain plane, calculated by taking the minimum value of the distances between the current frequency point coordinates and the coordinates of all frequency points on the boundary of the phase protection zone.
[0054] When filling in the preset offset, the shortest distance from each frequency point outside the phase protection zone boundary to the boundary is first calculated, and then this distance is mapped to the corresponding offset amplitude. The mapping function can be either linear or nonlinear. Under linear mapping, the offset amplitude is directly proportional to the shortest distance; under nonlinear mapping, a power function mapping can be used, so that the offset increases more slowly near the boundary and more quickly far from the boundary, achieving the effect of applying a more significant phase change to far-distance frequency points and a more gradual phase change to near-boundary frequency points. The sign of the offset can be uniformly set to positive or negative, or different signs can be assigned according to the azimuth angle of the current frequency point relative to the phase protection zone, so that the overall phase adjustment has directionality.
[0055] Step S245: Fill the frequency points inside the phase protection zone boundary with zero offset in the phase disturbance distribution, so that the phase disturbance distribution presents a zero value region inside and a gradually increasing offset transition zone outside.
[0056] Zero offset indicates that the original phase value of the frequency point is not changed when the phase indication distribution is applied; the frequency point maintains its inherent phase relationship in the defect-free panel. The coordinates of all frequency points corresponding to the boundary of the phase protection zone in the phase perturbation distribution matrix are filled with zeros. After filling, the phase perturbation distribution forms a two-layer structure on the frequency domain plane: an inner zero-value region and an outer offset transition band. The zero-value region is the frequency domain space occupied by one or more phase protection zones, where the value of any frequency point is zero. The offset transition band surrounds the zero-value region, and the offset values within it monotonically increase from the inside out according to the shortest distance, while the offset in the outer region far from the zero-value region reaches a preset upper limit value.
[0057] Step S246: The filled phase perturbation distribution is used as the phase cue distribution. The phase cue distribution is used to perform position-by-position combination processing with the phase spectrum representation of each image block when called.
[0058] The completed phase perturbation distribution is a two-dimensional real matrix with the same frequency domain size as the image patch phase spectrum. The value stored at each position is the offset correction applied to the phase at that frequency point when the phase cue distribution is invoked. This phase perturbation distribution is formally designated as the phase cue distribution and stored in the corresponding image patch group location buffer in system memory for subsequent phase enhancement processing of each image patch in the panel image to be detected. The position-by-position combination processing of the phase cue distribution and the image patch phase spectrum representation involves adding each element value in the phase cue distribution matrix to the corresponding phase value in the phase spectrum representation matrix position-by-position to obtain the adjusted phase value. This adjusted phase value is the value of the corresponding frequency point in the enhanced phase spectrum representation.
[0059] Step S250: Combine the amplitude spectrum representation of each image block with the amplitude cue distribution position by position. Based on the response amount less than the reference flux in the amplitude cue distribution, compress the amplitude amount at the set of main frequency points of the background texture. Based on the response amount of the reference flux in the amplitude cue distribution, retain the amplitude amount at other frequency points to obtain the enhanced amplitude spectrum representation.
[0060] Position-by-position combination processing multiplies the values of two elements in the same row and column indices of the amplitude spectrum representation matrix and the amplitude cue distribution matrix. The combination operation is defined as multiplying the amplitude value at a frequency point in the amplitude spectrum representation by the response at the corresponding frequency point in the amplitude cue distribution. The product is used as the amplitude value at that frequency point in the enhanced amplitude spectrum representation. When the frequency point is within the set of dominant frequency points of the background texture, the response at that frequency point in the amplitude cue distribution is less than the baseline flux. The product reduces the original amplitude value, achieving compression and suppression of the corresponding spectral energy of the background texture. The degree of compression is determined by the difference between the response and the baseline flux; the smaller the response, the stronger the compression. When the frequency point is outside the set of dominant frequency points of the background texture, the response at that frequency point in the amplitude cue distribution is equal to the baseline flux. The product maintains the original amplitude value, ensuring complete preservation of the amplitude at other frequency points. After performing this position-by-position multiplicative combination on all frequency points sequentially, the resulting new amplitude matrix is the enhanced amplitude spectrum representation of the image patch. In the enhanced amplitude spectrum representation, the frequency domain energy of the panel background texture is significantly compressed, while the frequency domain energy of non-periodic features such as defects is relatively prominent.
[0061] Step S260: Combine the phase spectrum representation of each image block with the phase cue distribution position by position, adjust the phase values at frequency points outside the phase stable frequency point set based on the preset offset in the phase cue distribution, and retain the phase values within the phase stable frequency point set to obtain the enhanced phase spectrum representation.
[0062] Position-by-position combination processing involves additively combining two elements at the same row and column indices in the phase spectrum representation matrix and the phase cue distribution matrix. The combination operation is defined as adding the original phase value at the frequency point in the phase spectrum representation to a preset offset at the corresponding frequency point in the phase cue distribution. The sum is used as the phase value at that frequency point in the enhanced phase spectrum representation. When the frequency point is within the set of phase-stable frequency points (i.e., within the phase protection zone boundary), the value at that frequency point in the phase cue distribution is filled with 0, and the additive combination does not change the original phase value, maintaining the phase relationship of that frequency point. When the frequency point is outside the set of phase-stable frequency points (i.e., within the offset transition zone), the value at that frequency point in the phase cue distribution is a non-zero preset offset. The additive combination causes a phase shift at that frequency point, the magnitude and direction of which are determined by the preset offset. After performing this position-by-position additive combination on all frequency points sequentially, the summation result needs to be phase periodically adjusted. This involves adjusting phase values that may exceed the -π to +π principal value range back to the principal value range through integer multiples of ±2π, resulting in the enhanced phase spectrum representation of the image patch. In the enhanced phase spectrum representation, the phase structure of stable regions is preserved intact, while the phase of unstable regions is artificially perturbed, and the changes in the phase mode of the defective region relative to the background are amplified.
[0063] Step S300: Perform inverse Fourier transform on the enhanced amplitude spectrum representation and enhanced phase spectrum representation of each image block to reconstruct the enhanced image block, and stitch all the enhanced image blocks together according to the division and arrangement order to obtain the complete enhanced image.
[0064] The inverse Fourier transform is an operation that converts the enhanced amplitude spectrum and enhanced phase spectrum, represented in the frequency domain, back into a spatial domain image. This operation resynthesizes the frequency components, after selective amplitude suppression and phase adjustment, into a pixel grayscale distribution in the spatial domain. An enhanced image block is the spatial domain result image obtained after performing an inverse Fourier transform on an image block. In this image, the display intensity of the panel background texture is suppressed, and the visual salience of defective areas is enhanced. The segmentation and arrangement order refers to the grid position index assigned to the image blocks during the segmentation process in step S100. Each image block has a corresponding horizontal and vertical arrangement number, and its spatial position in the complete enhanced image is determined according to these two numbers during stitching. The complete enhanced image is an enhanced image of the same size as the original panel image to be detected, obtained by placing all enhanced image blocks back into their corresponding spatial regions of the entire image according to their respective horizontal and vertical arrangement numbers, and then performing seam fusion processing.
[0065] In one embodiment, step S300 may specifically include the following steps S310 to S360: Step S310: Combine the enhanced amplitude spectrum representation and the enhanced phase spectrum representation of each image block into a complex spectrum representation. The real part of the complex spectrum representation is determined by the cosine components of the enhanced amplitude spectrum representation and the enhanced phase spectrum representation through a correspondence relationship, and the imaginary part is determined by the sine components of the enhanced amplitude spectrum representation and the enhanced phase spectrum representation through a correspondence relationship.
[0066] The complex spectrum representation is the input data structure for the inverse Fourier transform. It is a two-dimensional complex matrix with the same size as the enhanced amplitude spectrum representation. Each element in the matrix stores a complex value, which is determined in polar coordinates by the magnitude equal to the amplitude value of the corresponding frequency point in the enhanced amplitude spectrum representation and the argument equal to the phase value of the corresponding frequency point in the enhanced phase spectrum representation. When converting the polar coordinate expression to a Cartesian coordinate expression, the real part of the complex number is calculated as the product of the cosine of the amplitude value and the cosine of the phase value at that frequency point, and the imaginary part is calculated as the product of the amplitude value and the sine of the phase value at that frequency point. The real and imaginary parts of all frequencies are calculated one by one according to the above conversion rules to fill and form a complete complex spectrum representation matrix. Before performing the inverse Fourier transform, the complex spectrum representation matrix needs to undergo a quadrant decentering shift operation, that is, the zero frequency is moved from the center position back to the four corner positions. The shifting method is the inverse operation of the quadrant centering shift after the Fourier transform in step S100.
[0067] Step S320: Perform a two-dimensional discrete Fourier inverse transform on the complex spectral representation to convert the frequency domain complex information to the spatial domain, thereby obtaining the initial reconstructed image block of the image block.
[0068] The two-dimensional inverse discrete Fourier transform (IDFT) maps a frequency domain complex matrix back to a spatial domain real grayscale image. Its execution process is the inverse of the two-dimensional IFT: first, a one-dimensional IFT is performed on each column of the complex spectrum representation; then, a one-dimensional IFT is performed on each row of the column transformation result matrix, ultimately yielding a spatial domain real matrix of the same size as the image patch. The computational mechanism of the one-dimensional IFT is as follows: for a frequency domain complex sequence of length L, at each spatial index n, the k-th complex value in the frequency domain sequence is multiplied and accumulated with a rotation factor of angle +2π×k×n / L. All components from 0 to L-1 of k are accumulated and then divided by L for normalization, yielding the grayscale value at that spatial location. After performing a complete two-dimensional discrete Fourier inverse transform on the complex spectrum representation, the values of each element in the resulting real matrix may have a small imaginary part remaining due to the limited precision in the calculation process. In this case, the real part of each element is taken and rounded to a reasonable gray range to obtain the initial reconstructed image block of the image block.
[0069] In one embodiment, step S320, performing a two-dimensional discrete Fourier inverse transform on the complex spectral representation to convert the frequency domain complex information to the spatial domain, and obtaining the initial reconstructed image patch of the image patch, may specifically include the following steps S321 to S326: Step S321: Perform a two-dimensional discrete Fourier inverse transform on the complex spectral representation to obtain the spatial domain real matrix of the image patch as the reconstruction image to be processed.
[0070] The reconstructed image to be processed is a preliminary spatial domain reconstructed image obtained by direct inverse transformation after the frequency domain amplitude and phase have been adjusted. Its pixel size is consistent with the original image patch. In this reconstructed image to be processed, the intensity of the background texture has been initially weakened, but the contrast between the defective area and the background texture preservation area may not have reached the optimal detection state. Further region-adaptive contrast optimization is needed in subsequent steps.
[0071] Step S322: Obtain the response difference of each frequency point in the amplitude indication distribution relative to the reference flux, arrange the response difference along the frequency domain coordinates to form a frequency domain suppression depth record, and perform an inverse Fourier transform on the frequency domain suppression depth record to obtain a spatial domain suppression intensity distribution map. The pixel positions of the spatial domain suppression intensity distribution map correspond one-to-one with the pixel positions of the reconstruction map to be processed.
[0072] The response difference is the difference between the response value at each frequency point in the amplitude indication distribution and the reference flux. This difference reflects the depth of spectral energy suppression at that frequency point: a positive and larger difference indicates deeper suppression at that frequency point. The frequency domain suppression depth record is a real matrix arranged according to the frequency domain coordinates of all frequency points. Its size is exactly the same as the amplitude indication distribution. This matrix records the energy reduction distribution of each frequency component in the frequency domain plane. The process of performing an inverse Fourier transform on the frequency domain suppression depth record is similar to the process of performing an inverse transform on the complex spectrum in step S321. The difference is that the input matrix is a real matrix instead of a complex matrix. Therefore, it is regarded as a complex matrix with all imaginary parts being zero. It is processed according to the same two-dimensional discrete inverse Fourier transform algorithm to obtain the spatial domain suppression intensity distribution map. The spatial domain suppression intensity distribution map is a real matrix with the same size as the image to be reconstructed. The value of each pixel position in the map represents the degree of intensity attenuation caused by the frequency domain suppression operation at that spatial position. The deeper the suppression, the higher the pixel value. The pixel-by-pixel correspondence between the spatial domain suppression intensity distribution map and the reconstructed map to be processed is guaranteed by the linearity and shift invariance of the inverse Fourier transform. The effect of the suppression operation at a specific position in the frequency domain is faithfully reflected in the corresponding pixel position in the spatial domain.
[0073] Step S323: In the spatial domain suppression intensity distribution map, the continuous areas with suppression intensity lower than the preset intensity limit are defined as texture preservation zones, and the continuous areas with suppression intensity not lower than the preset intensity limit are defined as defect highlighting zones.
[0074] Texture preservation regions are spatially continuous areas with relatively low suppression levels in the spatial domain suppression intensity distribution map. These regions correspond to areas where the background texture is significantly weakened after frequency domain modulation but still exists, mainly containing the normal texture components of the panel. Defect highlighting regions are spatially continuous areas with relatively high suppression levels in the spatial domain suppression intensity distribution map. These regions correspond to areas where the background texture is deeply suppressed after frequency domain modulation, and the frequency domain response caused by defects is highlighted, with a high probability of defect presence. A preset intensity limit is a suppression intensity threshold used to divide the texture preservation and defect highlighting regions. This threshold is determined by calculating the statistical distribution of all pixel values in the spatial domain suppression intensity distribution map, taking a quantile value of this distribution as the boundary between low and high suppression areas, ensuring that the texture preservation region can cover most of the normal background texture areas, while the defect highlighting region can capture abnormal areas with significantly high suppression intensity. Connected component analysis is used to extract continuous regions. The spatial domain suppression intensity distribution map is binarized using the preset intensity limit, and then connected component labeling is performed on the binary image to obtain spatial region masks for the texture preservation and defect highlighting regions, respectively.
[0075] Step S324: Based on the boundaries of the texture preservation partition and the defect highlighting partition, construct a binarized region segmentation template. The binarized region segmentation template takes the first marker value at the texture preservation partition position and the second marker value at the defect highlighting partition position.
[0076] The binarized region segmentation template is a binary matrix of the same size as the reconstructed image to be processed. Each pixel in the matrix stores a region classification label. The first label value is distinct from the second label value, and they can take any different numerical pairs, for example, the first label value is 0 and the second label value is 1. When constructing the template, all pixel coordinates of the reconstructed image to be processed are traversed. If the coordinate is classified as belonging to the texture preservation zone in step S323, the template value at that position is assigned the first label value; if the coordinate is classified as belonging to the defect highlighting zone, the template value at that position is assigned the second label value. The spatial distribution of the two label values in the template completely depicts the geometric shape and spatial range of the two types of regions with different properties in the reconstructed image to be processed.
[0077] Step S325: Using a binarized region segmentation template, the contrast of the pixel set located in the defect highlighting region in the image to be reconstructed is expanded to increase the grayscale range of the pixel set. At the same time, the grayscale distribution of the pixel set located in the texture preservation region is preserved so that the grayscale statistical characteristics of the pixel set are consistent with those before processing.
[0078] Contrast expansion applies a stretching transformation to the pixel grayscale values within the defect highlighting zone, linearly mapping the originally concentrated grayscale range to a wider grayscale interval. Specifically, it calculates the minimum and maximum grayscale values of all pixels within the current defect highlighting zone, maps the minimum value to the lower limit of the target grayscale range, maps the maximum value to the upper limit of the target grayscale range, and maps intermediate values using linear interpolation. Setting the lower and upper limits of the target grayscale range increases the range of pixel grayscale values within the defect highlighting zone, thereby enhancing the discernibility of details within the defect area and the contrast between the defect and the surrounding background. Grayscale distribution preservation applies no transformation or only a translation operation that maintains the relative grayscale relationships to the pixel grayscale values within the texture preservation zone. This ensures that the statistical characteristics of the pixels in this region, such as the shape of the grayscale histogram, the mean grayscale value, and the variance grayscale value, remain essentially consistent before and after processing, maintaining the natural visual appearance of normal texture areas. The process of processing the reconstructed image to be processed is as follows: traverse the reconstructed image to be processed pixel by pixel, and determine whether the pixel accepts contrast expansion based on the marker value of the corresponding position in the binarized region division template: pixels with a marker value of the second marker value perform contrast expansion mapping, while pixels with a marker value of the first marker value keep their original grayscale value unchanged.
[0079] Step S326: Merge the defect highlighting partition pixels after contrast expansion and the texture preservation partition pixels after grayscale distribution preservation according to their spatial positions to obtain the initial reconstructed image block of the image block.
[0080] The merging operation combines two parts of pixels processed in different ways into a complete image based on their original pixel coordinates. For all pixel coordinates in the image to be reconstructed, if the coordinate belongs to a defect-highlighting zone, the pixel grayscale value at that location is taken from the contrast-expanded zone result and filled into the corresponding position in the merged image; if the coordinate belongs to a texture-preserving zone, the pixel grayscale value at that location is taken from the grayscale-preserved zone result and filled into the corresponding position in the merged image. After the entire image is traversed and filled, the resulting image is the initial reconstructed image patch with region-adaptive contrast optimization. In this initial reconstructed image patch, the visual representation of the background texture area remains consistent with the original image to be reconstructed, while the grayscale range of the defect area is significantly expanded, enhancing the visual contrast between the defect and the background.
[0081] Step S330: Perform dynamic range stretching on the initial reconstructed image patch to map the pixel value distribution to a preset visible range, maintain the relative contrast level between the defective area and the background, and generate an enhanced image patch.
[0082] Dynamic range stretching (VRLT) is a process that linearly scales and translates the pixel grayscale values of an initially reconstructed image patch to a common image display or model input range while maintaining their relative sizes. The preset visible range is the standard grayscale value range used for image storage or display, such as the integer range from 0 to 255. The VRLT process involves: calculating the minimum and maximum grayscale values of all pixels in the initial reconstructed image patch; constructing a linear mapping function; taking the initial grayscale values as input and outputting the scaled and translated grayscale values; and ensuring that the initial minimum corresponds to the lower limit of the preset visible range and the initial maximum corresponds to the upper limit. This linear mapping function is applied pixel-by-pixel to each grayscale value of the initial reconstructed image patch. The coefficients and bias terms in the mapping process remain consistent across all pixels, thus preserving the relative contrast between defective areas and the background without distorting the relative grayscale relationship due to stretching. The resulting image after mapping is the enhanced image patch of that patch.
[0083] Step S340: Determine the placement coordinates of each enhanced image block in the complete image according to the horizontal and vertical arrangement numbers of the enhanced image blocks in the panel image to be detected.
[0084] In step S100, when dividing the image blocks, each image block is assigned a horizontal and vertical arrangement number. The horizontal arrangement number records the column number of the image block from left to right in the panel image to be detected, and the vertical arrangement number records the row number of the image block from top to bottom in the panel image to be detected. The standard block size of the image block is a known fixed value during division. Therefore, the placement coordinates of each enhanced image block in the complete enhanced image can be directly derived from the horizontal and vertical arrangement numbers: the starting horizontal coordinate is equal to the horizontal arrangement number multiplied by the standard block width, and the starting vertical coordinate is equal to the vertical arrangement number multiplied by the standard block height. The placement coordinates determine the rectangular area occupied by each enhanced image block in the complete enhanced image. By iteratively performing the placement coordinate calculation on all enhanced image blocks, a spatial positioning framework can be laid out for the stitching operation.
[0085] Step S350: Perform transition fusion processing at the stitching boundary between adjacent enhanced image blocks. Based on the pixel change gradient on both sides of the boundary between two adjacent enhanced image blocks, generate a gradient blending weight to make the pixel values at the stitching seam transition smoothly.
[0086] The stitching boundary is the straight line connecting two horizontally or vertically adjacent enhanced image patches in a complete enhanced image. The transition blending process, within a narrow blending transition zone on both sides of the stitching boundary, does not directly use the original boundary pixel values of each enhanced image patch. Instead, it weights and mixes the pixel grayscale values of the two enhanced image patches within this region, allowing the pixel grayscale values within the blending transition zone to naturally transition from one image patch to the other. The pixel change gradient is the rate of change of pixel grayscale values of an enhanced image patch in a direction perpendicular to the boundary. A large gradient indicates a drastic grayscale change near the boundary on that side, while a small gradient indicates a gradual grayscale change near the boundary. Generating gradient blending weights based on the pixel change gradient involves adaptively adjusting the shape and rate of change of the weight function within the blending transition zone with reference to the pixel change gradients on both sides. This ensures that the weights on the side with drastic grayscale changes quickly give way to the weights on the side with gradual grayscale changes, while the weights on the side with gradual grayscale changes transition slowly, so that the final grayscale distribution at the seam conforms to the original trend of the two images.
[0087] In one embodiment, step S350 may specifically include the following steps S351 to S356: Step S351: Extract the inner pixel strips of the first enhanced image block and the second enhanced image block adjacent to the current stitching boundary. The width of the two pixel strips corresponds to the set half-width of the fusion transition area.
[0088] The first and second enhanced image blocks are two adjacent enhanced image blocks sharing the same stitching boundary. The inner pixel stripe is a rectangular pixel strip region formed by extending a predetermined half-width distance into the enhanced image block, based on the stitching boundary. Pixels within this region will be used for subsequent brightness gradient analysis and blending processing. The predetermined half-width of the blending transition region is a spatial distance parameter pre-set based on the standard block size and texture scale of the image blocks. The total width of the blending transition region is twice the predetermined half-width, spanning from the predetermined half-width inside the boundary of the first enhanced image block to the predetermined half-width inside the boundary of the second enhanced image block. The extraction operation copies all pixels in the first enhanced image block along the vertical boundary direction inwards along the predetermined half-width direction to form the first pixel stripe, and copies all pixels in the second enhanced image block along the vertical boundary direction inwards along the predetermined half-width direction to form the second pixel stripe.
[0089] Step S352: Determine the brightness gradient distribution of the first enhanced image block pixel strip along the direction perpendicular to the boundary, and the brightness gradient distribution of the second enhanced image block pixel strip along the direction perpendicular to the boundary.
[0090] The brightness gradient is the rate of change of a pixel's grayscale value along a specified spatial direction. For the pixel strip of the first enhanced image block, the first-order difference of grayscale values is calculated pixel by pixel along a direction perpendicular to the boundary. The difference is calculated as the absolute value of the difference between the grayscale values of two adjacent pixels along the perpendicular boundary direction. Arranging all the difference values according to their positions yields the brightness gradient distribution on the first enhanced image block side. The same first-order difference calculation is performed on the pixel strip of the second enhanced image block in the same direction and manner to obtain the brightness gradient distribution on the second enhanced image block side. During the gradient calculation process, the gradient of the outermost pixel at the boundary of the strip adopts the same gradient value as the immediately adjacent inner pixel to ensure the continuity of the gradient distribution at the boundary.
[0091] Step S353: Normalize and align the brightness gradient distributions on both sides to ensure that the gradient change rates of the two enhanced image blocks at the boundary are consistent, and generate a brightness gradient matching curve.
[0092] Normalization alignment involves dividing the brightness gradient distributions on both sides by the maximum value of their respective gradient distributions, scaling the gradient values to a common numerical space between 0 and 1. This eliminates the incomparability of absolute gradient values caused by differences in absolute grayscale levels between the two images, allowing the gradient change rates to be mutually referenced in a relative sense. After normalization, the normalized gradient distribution on the first enhanced image block side is extended outward along the vertical boundary at the boundary, and the normalized gradient distribution on the second enhanced image block side is extended outward in the opposite direction along the vertical boundary. The two are then combined at the boundary to generate a brightness gradient matching curve spanning the entire fusion transition zone. The segment of this curve near the first enhanced image block reflects the gradient change rate of the first enhanced image block, and the segment near the second enhanced image block reflects the gradient change rate of the second enhanced image block. At the middle boundary, the values on both sides of the curve tend to be continuous after normalization.
[0093] Step S354: Construct a gradient weight function based on the brightness gradient matching curve. The gradient weight function assigns the weight of the first enhanced image block as the dominant weight at the beginning of the fusion transition zone and the weight of the second enhanced image block as the dominant weight at the end. The middle position transitions according to the gradient rule.
[0094] The gradient weighting function defines the proportion of the pixel values of the first and second enhanced image blocks at each pixel location within the blending transition region in the mixed output. The weighting function is designed based on the brightness gradient matching curve: in sections where the gradient matching curve rises rapidly, the weight change rate increases accordingly; in sections where the gradient matching curve rises slowly, the weight change rate decreases accordingly. Specifically, the brightness gradient matching curve is accumulated and summed along the blending transition region from the starting point on the first enhanced image block side to obtain the accumulated gradient curve. This accumulated gradient curve is then normalized by dividing it by the total accumulated value of the accumulated gradient curve. The resulting normalized accumulated gradient curve is the weighting function assigned to the second enhanced image block. The weighting function assigned to the first enhanced image block is a constant value of 1 minus the weighting function assigned to the second enhanced image block. At the beginning of the fusion transition zone, the weight of the second enhanced image block is 0 and the weight of the first enhanced image block is 1, and the mixed output is completely dominated by the first enhanced image block; at the end, the weight of the second enhanced image block is 1 and the weight of the first enhanced image block is 0, and the mixed output is completely dominated by the second enhanced image block; at any intermediate position, the sum of the two weights is always 1, and the weights change smoothly and gradually along the direction of the fusion transition zone.
[0095] Step S355: Based on the gradient weighting function, perform a blending process on each pixel of the first enhanced image block and the second enhanced image block in the fusion transition zone to generate a sequence of fusion boundary pixel values.
[0096] The blending process involves, for each pixel location within the fusion transition region, acquiring the pixel grayscale value of the first enhanced image block and the second enhanced image block at that location, along with the corresponding weights of the first and second enhanced image blocks. A weighted sum is then calculated: the first grayscale value multiplied by the first weight plus the second grayscale value multiplied by the second weight. The result is used as the fusion boundary pixel value for that location. This blending process is applied to all pixel locations within the fusion transition region to generate a complete sequence of fusion boundary pixel values.
[0097] Step S356: Replace the original pixels at the corresponding stitching boundary with the pixel value sequence of the fusion boundary, so that the first enhanced image block and the second enhanced image block form a continuous image region along the stitching boundary, thus forming an enhanced image block combination after eliminating block effects.
[0098] The replacement operation involves writing each pixel value from the fusion boundary pixel value sequence to the corresponding pixel position in the complete enhanced image, overwriting the original pixel value of the first or second enhanced image block stored at that position. After performing transition fusion processing and pixel replacement on each stitching boundary between all adjacent enhanced image blocks according to steps S351 to S355, the block boundaries between all enhanced image blocks are eliminated, and the enhanced image blocks are combined into a continuous image region with a spatially continuous transition in pixel grayscale, without any seams or artificial traces.
[0099] Step S360: Combine all the enhanced image blocks that have undergone the splicing boundary transition fusion process into a complete enhanced image. The size of the complete enhanced image is the same as that of the panel image to be detected.
[0100] The combined image of the enhanced image blocks after completing all boundary fusion and replacement operations in step S350 is output as a whole, thus obtaining the complete enhanced image. Since the placement coordinates and size of each enhanced image block are obtained based on the precise grid configuration when dividing the panel image to be detected, and the boundary fusion transition processing does not change the external size and spatial occupancy of the image blocks, the number of pixel rows and pixel columns of the combined complete enhanced image are completely consistent with the original panel image to be detected, and the two are strictly equivalent in geometric dimensions.
[0101] Step S400: Input the fully enhanced image into a pre-trained detection model with fixed response characteristics. The pre-trained detection model performs defect localization and category prediction on the fully enhanced image and outputs the current detection result, which includes defect location markers and defect category markers.
[0102] The pre-trained detection model is a deep convolutional neural network model that has been fully trained and whose weights are not updated during detection applications. All convolutional kernel weights, bias terms, scaling and offset coefficients of batch normalization layers, and weights of fully connected layers are fixed, ensuring the same output regardless of the input image. Fixed response characteristics mean that the values of all learnable parameters are not modified during inference, guaranteeing repeatability and stability of the detection results. Defect localization determines the precise spatial extent of each defect in the enhanced image, typically represented by a rectangular bounding box. Category prediction determines which specific category in a predefined defect category set each located defect region belongs to. The predefined defect category set is a list of defect types predefined according to the quality inspection standards for panel products, including categories such as scratches, dents, dirt, cracks, and bubbles. The current detection results consist of defect location markers and defect category markers: the defect location markers record the location information of all detected defect bounding boxes using a set of coordinate sequences, and each bounding box can be uniquely determined by the coordinates of its upper left corner and lower right corner in the image; the defect category markers give the defect type of each bounding box using a sequence of category labels that correspond one-to-one with the bounding box.
[0103] In one embodiment, step S400 may specifically include the following steps S410 to S460: Step S410: The complete enhanced image is fed into the deep convolution processing stream of the pre-trained detection model. The image content is extracted step by step through multiple convolution processing layers configured in series to generate a multi-level abstract representation map set. The receptive field coverage of each abstract representation map increases as the layer deepens.
[0104] A deep convolutional processing flow is a hierarchical feature extraction pathway consisting of multiple convolutional processing layers strung together. The fully enhanced image enters from the input and is sequentially fed to each convolutional processing layer. Each convolutional processing layer receives the feature representation map from the previous layer, performs its own internal operations, and outputs a further processed and refined new feature representation map. This output is simultaneously passed to the next layer and stored in a bypass path. An abstract representation map is a multi-channel feature matrix output from a convolutional processing layer in the convolutional processing flow. Each channel captures a specific visual pattern or semantic attribute of the image. A multi-level abstract representation map set is a sequence of representation maps formed by summarizing the outputs of all convolutional processing layers in order of processing depth. Shallow abstract representation maps have high spatial resolution and small receptive fields, recording the geometric details and local texture information of the image; deep abstract representation maps have low spatial resolution and large receptive field coverage, recording the global semantic layout and overall structural context of the image.
[0105] In one embodiment, step S410 may specifically include the following steps S411 to S416: Step S411: Obtain the pre-fixed arrangement order of convolutional processing layers in the pre-trained detection model. The convolutional processing layers include alternately set filter expansion layers and compression aggregation layers. The increase of filter expansion layers represents the number of channels, and the reduction of compression aggregation layers represents the plane size.
[0106] The pre-trained detection model's main structure employs a deep convolutional neural network architecture with recognized performance in image detection, such as ResNet deep residual network as the backbone feature extraction network. Several stages of this network correspond to the arrangement of convolutional processing layers. The filtering expansion layer corresponds to the convolutional segment in the backbone network that increases the number of feature channels. This type of layer contains batch normalization layers and modified linear unit activation functions, expanding the number of representation channels by increasing the number of convolutional kernels, thus enhancing the expressive power of the feature space dimension. The compression and aggregation layer corresponds to the downsampling segment in the backbone network that reduces the planar size of the feature map. This type of layer reduces the spatial height and width of the output feature map by using convolutional operations with a stride greater than one or by using pooling operations, reducing planar resolution while expanding the receptive field coverage of each feature location. The filtering expansion layer and the compression and aggregation layer alternate in the deep convolutional processing flow, forming a pyramid-shaped feature transformation skeleton with increasing representation channel number and decreasing planar size between layers.
[0107] Step S412: Input the complete enhanced image into the first convolutional processing layer, perform filtering expansion and nonlinear activation to obtain an initial level abstract representation map. The initial level abstract representation map has the expanded number of channels and retains a planar size close to that of the original image.
[0108] The first convolutional processing layer serves as the entry point for the entire depthwise convolutional processing flow. It contains a large set of convolutional filters with small spatial dimensions and a stride of 1, ensuring the convolution operation maintains a relatively constant spatial resolution of the input fully enhanced image. The fully enhanced image serves as input, with the number of channels corresponding to the single-channel grayscale image. After entering the first convolutional processing layer from the input layer, it undergoes two-dimensional cross-correlation with multiple sets of convolutional kernels. The result is a multi-channel output equal to the number of convolutional kernels, expanding the number of channels compared to the single-channel number of the input image. Immediately after the cross-correlation operation, the feature distribution is standardized by a batch normalization layer, and then non-linear activation is performed using a modified linear unit activation function, generating an initial-level abstract representation. This initial-level abstract representation retains a height and width close to the fully enhanced image, but the channel dimension has been expanded from a single channel to the preset number of channels.
[0109] Step S413: Pass the current level abstract representation graph to the next convolutional processing level, perform compression aggregation to reduce the plane size, and perform further filtering expansion to increase the number of channels to obtain the next level abstract representation graph. Repeat this passing process until all convolutional processing levels are traversed.
[0110] Each level of abstract representation map undergoes a compression-aggregation operation before being passed to the next convolutional processing layer. This compression-aggregation is achieved using a convolutional operation with a stride of a preset compression ratio. During the sliding computation, this operation skips certain pixel positions with a specified stride, reducing the spatial size of the output feature map to one-half of the preset compression ratio of the input, while maintaining or moderately varying the number of channels. Subsequent filtering and expanding operations further increase the number of channels by increasing the output channels of the convolutional kernel. The output after merging compression-aggregation and filtering / expanding is used as the next level of abstract representation map. Compared to the current level's abstract representation map, this map has a smaller planar size, more channels, and a wider receptive field. Following the order of the convolutional processing layers, the output of this level is used as the input for the next level, cyclically progressing until the last convolutional processing layer completes its computation.
[0111] Step S414: After each compression aggregation level, the planar size of the abstract representation map output by the compression aggregation level becomes the preset compression ratio of the planar size of the input abstract representation map. At the same time, the image structure granularity captured by each channel of the output abstract representation map is coarser than that of the input.
[0112] The preset compression ratio is determined by the stride parameter of the convolution operation in the compression aggregation level. The stride value is directly equal to the one-dimensional size reduction factor of the compression ratio. When the stride is greater than one, each spatial location of the output feature map is calculated from the set of pixels with a fixed sampling interval in the input feature map. The input spatial range integrated at each output location is expanded, thus the spatial range covered by the image structure captured by each channel becomes larger, and the geometric granularity of the feature representation tends to coarsen. After multiple levels of compression aggregation, the spatial resolution of the abstract representation map gradually decreases, and the area of the original image corresponding to a single element in the abstract representation map, i.e., the receptive field coverage, expands step by step.
[0113] Step S415: Record the intermediate abstract representation map output of each convolution processing layer to obtain a sequence of abstract representation maps sorted by processing depth. In the sequence of abstract representation maps, shallow abstract representation maps retain fine-grained geometric details, while deep abstract representation maps capture a wide range of semantic layouts.
[0114] When each convolutional processing layer outputs its own abstract representation map, a copy of the output representation map is bypassed from the depthwise convolutional processing stream and stored in a memory abstract representation map buffer. The buffer stores all outputs from shallow to deep layers in sequence. After being arranged in the output order, the resulting sequence of abstract representation maps exhibits a hierarchical evolution of features: shallow-level representation maps have large planar dimensions and few channels, with the activation values of each channel mainly responding to low-level visual features such as edge direction, texture granularity, and color contrast; deep-level representation maps have significantly smaller planar dimensions and many channels, with the activation values of each channel mainly responding to high-level semantic features such as part shape, target contour, and contextual regions.
[0115] Step S416: The sequence of abstract representation graphs is used as a multi-level abstract representation graph set. The spatial resolution and channel dimension of each abstract representation graph in the multi-level abstract representation graph set are distributed in a pyramid shape from shallow to deep.
[0116] The multi-level abstract representation atlas is a logical collection of all abstract representation maps sequentially saved in step S415. Because the planar size decreases and the number of channels increases progressively from shallow to deep levels, this atlas has a pyramidal structure in spatial dimensions. Subsequent operations of the pre-trained detection model will extract representation maps of different levels from this pyramidal atlas for region localization and category determination.
[0117] Step S420: Extract the deepest level of the global semantic abstract representation graph from the multi-level abstract representation graph set, perform region activation mapping on the global semantic abstract representation graph, and locate suspected defective activation regions whose response intensity exceeds the preset response threshold.
[0118] The deepest-level global semantic abstraction map (GSAD) sits at the top of the abstraction map pyramid, possessing the smallest spatial resolution and the most channels. Each channel's activation map globally describes the distribution of specific high-level semantic patterns in the complete augmented image. Region activation mapping (REM) processes this GSAD to generate a two-dimensional activation intensity map corresponding to the spatial location in the complete augmented image. The value at each location in the map represents the overall confidence that any type of defect exists at that location. REM is implemented by inputting the GSAD into a detection head module consisting of a single convolutional layer and a non-linear activation function. The convolutional layer outputs one channel. The convolutional layer weights and compresses the activation values of all channels into a single-channel defect activation intensity map. This is then compressed to between 0 and 1 using a sigmoid non-linear activation function, resulting in the two-dimensional activation intensity map. The predicted response threshold is a pre-defined activation intensity discrimination threshold. Each spatial location in the activation intensity map is traversed, and locations with activation intensities exceeding this threshold are marked as foreground, while the rest are marked as background. Adjacent foreground locations in space are aggregated into a connected region, and each connected region is a potential defect activation region. All potential defect activation regions are then combined into the current candidate region set. This localization process generates defect candidate regions regardless of category; it does not distinguish between defect types and only identifies the approximate spatial location of potential defects.
[0119] Step S430: Select at least two shallow level detail abstraction maps from the multi-level abstraction map set, align the shallow level detail abstraction maps with the suspected defect activation areas in space, and crop out the local detail representations corresponding to each suspected defect activation area.
[0120] Shallow-level detail abstraction maps are selected from the shallower levels of the abstraction map pyramid. These maps possess high spatial resolution and rich geometric detail information, providing ample feature support for accurate boundary localization and fine-grained category determination of suspected defective activation regions. Selecting at least two shallow levels ensures feature complementarity across different spatial resolutions and receptive field coverages. Spatial alignment is performed by mapping the spatial coordinates of suspected defective activation regions identified in the deep-level global semantic abstraction map to the corresponding spatial window in the shallow-level representation map, based on the spatial resolution ratio between the deep-level global semantic abstraction map and the shallow-level detail abstraction map. Multi-channel local feature regions are then cropped from the shallow-level representation map according to this window. A corresponding local detail representation is cropped from each suspected defective activation region in different shallow-level representation maps. These local detail representations are then fed as multi-scale features into subsequent bounding box regression and category prediction processing.
[0121] Step S440: Perform bounding box offset regression on the local detail representation of each suspected defect activation area to generate a corrected precise location bounding box. The precise location bounding box indicates the spatial extent of the defect using the coordinates of the upper left corner and the lower right corner.
[0122] Bounding box offset regression performs fine-tuning on the position and size of the initial coarse region given by the suspected defect activation area, ensuring that the bounding box accurately fits the actual boundary of the defect. Offset regression is performed by a dedicated regression branch network in the pre-trained detection model. This branch network takes local detail representations as input and outputs center position correction parameters and span correction parameters corresponding to the current region. These correction parameters are then used to adjust the original region box to obtain the precise location bounding box.
[0123] In one embodiment, step S440 may specifically include the following steps S441 to S445: Step S441: For each suspected defective activation region, extract the center region from the corresponding local detail representation as the regression input region. Input the regression input region into the bounding box offset regression branch. The bounding box offset regression branch performs convolution compression processing on the regression input region and outputs the center correction component and the span correction component. The center correction component includes the offset description along the first coordinate axis and the offset description along the second coordinate axis. The span correction component includes the scaling description along the first coordinate axis and the scaling description along the second coordinate axis.
[0124] The bounding box offset regression branch consists of a small convolutional subnetwork containing several convolutional layers with small kernel sizes. Modified linear unit activation functions are used between layers, and a global average pooling layer is applied at the end to flatten the spatial dimensions. A fully connected layer then outputs four regression descriptors. The regression input region is a fixed-size feature block clipped from the local detail representation, with the center of the currently suspected defective activation region as its geometric center. The clipping size matches the expected input size of the bounding box offset regression branch, and any insufficient areas are filled with zero values. After being fed into the bounding box offset regression branch, the regression input region undergoes convolutional compression and layer operations. The fully connected layer outputs four real-valued descriptors. Two of these descriptors are center correction components, recording the center position offset along the first coordinate axis (horizontal) and the second coordinate axis (vertical), respectively. The other two are span correction components, recording the size scaling along the first and second coordinate axes, respectively.
[0125] Step S442: Based on the offset descriptor along the first coordinate axis and the offset descriptor along the second coordinate axis in the center correction component, adjust the original center position of the current suspected defect activation area to obtain the corrected center position. At the same time, based on the scaling descriptor along the first coordinate axis and the scaling descriptor along the second coordinate axis in the span correction component, adjust the original width and original height of the current suspected defect activation area to obtain the corrected area width and area height.
[0126] The original center position is the geometric center coordinate of the suspected defect activation region in the current coordinate reference frame. Its first coordinate axis coordinate is equal to half the sum of the coordinates of the left and right boundaries of the original region, and its second coordinate axis coordinate is equal to half the sum of the coordinates of the upper and lower boundaries of the original region. The position is adjusted by multiplying the center offset descriptor by the original region span in the corresponding direction, and then superimposing this onto the original center coordinates to obtain the corrected center coordinates. The original width and original height are equal to the difference between the right and left boundaries and the difference between the lower and upper boundaries of the current region, respectively. The size is adjusted by performing an exponential operation on the span scaling descriptor to obtain a scaling factor, then multiplying the original width by the scaling factor in the first coordinate axis direction to obtain the corrected region width, and multiplying the original height by the scaling factor in the second coordinate axis direction to obtain the corrected region height.
[0127] Step S443: Based on the corrected center position, the corrected area width, and the corrected area height, calculate the coordinates of the two sets of diagonal corner points of the rectangle surrounding the defect area. Record the corner point with the smaller value along the first coordinate axis as the first corner point coordinate, and the corner point with the larger value along the first coordinate axis as the second corner point coordinate.
[0128] The first corner coordinates include coordinate values along the first coordinate axis and coordinate values along the second coordinate axis. The coordinate value along the first coordinate axis is the first coordinate of the corrected center coordinate minus half the width of the corrected region. The coordinate value along the second coordinate axis is the second coordinate of the corrected center coordinate minus half the height of the corrected region. The second corner coordinates also include coordinate values along the first and second coordinate axes. The coordinate value along the first coordinate axis is the first coordinate of the corrected center coordinate plus half the width of the corrected region. The coordinate value along the second coordinate axis is the second coordinate of the corrected center coordinate plus half the height of the corrected region. The first and second corner coordinates together define the spatial extent of the blemish's rectangular bounding box in the two-dimensional plane of the image.
[0129] Step S444: Perform boundary compliance processing on the coordinates of the first corner point and the second corner point. Determine whether the rectangle formed by the coordinates of the first corner point and the second corner point is completely inside the corresponding image block. If there is a part that exceeds the boundary of the image block, cut the rectangle to the boundary of the image block along the corresponding boundary direction to obtain a compliant rectangle.
[0130] For boundary compliance, for example, the larger of the first coordinate axis direction value and the larger of the second coordinate axis direction value compared to zero is taken as the compliant first corner point; the smaller of the smaller of the first coordinate axis direction value compared to the image block width minus 1 and the smaller of the smaller of the second coordinate axis direction value compared to the image block height minus 1 is taken as the compliant second corner point. The rectangle enclosed by the compliant first and compliant second corner points is the truncated compliant rectangle. This compliant rectangle ensures that it does not exceed the effective pixel range of the corresponding image block.
[0131] Step S445: Use the coordinates of the first and second corner points of the compliant rectangle as the corner coordinates of the precise location bounding box corresponding to the suspected defect activation area. The corner coordinates of the precise location bounding box are used to form the defect location marker in the current detection result.
[0132] The precise location bounding box is the compliant rectangle, which is fully represented by the coordinates of the first and second compliant corner points. A unique bounding box identifier is assigned to this bounding box, and the coordinate data is organized into a record marking the defect location. The record includes the identifier, the lower limit of the first coordinate axis, the lower limit of the second coordinate axis, the upper limit of the first coordinate axis, and the upper limit of the second coordinate axis, to be summarized later.
[0133] Step S450: Perform category probability mapping on the local detail representation within each precise location bounding box, map the local detail representation to the probability distribution on the preset defect category set, and select the defect category corresponding to the highest probability in the probability distribution as the defect category label of the bounding box.
[0134] The category probability mapping is performed by a dedicated classification branch network in the pre-trained detection model. The classification branch network is structured as follows: the local detail representation of the region corresponding to the precise location bounding box is input into a classification sub-network containing several convolutional layers and one fully connected layer. The number of output nodes in the fully connected layer equals the total number of categories in the preset defect category set. After the fully connected layer output value is processed by a flexible maximum exponential normalization function, the probability distribution of the region across all preset defect categories is obtained. The flexible maximum function exponentializes the original score of each output node and divides it by the sum of the exponentialized scores of all nodes, ensuring that the output probability is non-negative and the sum is 1. The defect category label corresponding to the category with the largest value in the probability distribution is selected as the defect category label for that precise location bounding box.
[0135] Step S460: Collect the coordinates of all precise location bounding boxes and their corresponding defect category labels into the current detection result. The defect location labels in the current detection result are composed of the coordinate sequence of the precise location bounding boxes, and the defect category labels are composed of the defect category labels corresponding to each bounding box.
[0136] The aggregation process involves iterating through all suspected defect activation regions and generating each precise location bounding box. The corner coordinates and category label of each bounding box constitute a detection entry. All detection entries are aggregated to form the complete data volume of the current detection result. If the precise location bounding box obtained after bounding box offset regression for a suspected defect activation region is too small, it can be removed using a preset minimum size filtering rule and not included in the current detection result. The current detection result is output in a structured data format. The defect location label is a list, and each element in the list contains a bounding box coordinate quadruple. The defect category label is a list of category labels of the same length as the list above, with each list element corresponding one-to-one with a bounding box.
[0137] Step S500: Construct a frequency domain prompt adjustment amount based on the current detection result, and correct the amplitude prompt distribution and phase prompt distribution through the frequency domain prompt adjustment amount to obtain the corrected amplitude prompt distribution and the corrected phase prompt distribution. Then, based on the corrected amplitude prompt distribution and the corrected phase prompt distribution, re-execute the process of calling the pre-constructed amplitude prompt distribution and phase prompt distribution to output the current detection result until the prompt adjustment termination condition is met. Use the defect location mark and defect category mark of the last output as the final defect detection information.
[0138] The frequency domain cue adjustment is an error feedback signal that directionally corrects the amplitude and phase cue distributions based on the current detection results. This signal is calculated based on the difference between the detected defect areas and the ideal defect-free frequency domain representation. By superimposing the frequency domain cue adjustment onto the currently used amplitude and phase cue distributions, a gradient correction can be performed on the cue distribution, resulting in more significant frequency domain enhancement of defect areas or more effective background suppression in the next iteration. The cue adjustment termination condition is the convergence criterion for determining when the iterative optimization process can stop. When the change in the spatial overlap of defect location markers between two consecutive iterations is less than a preset convergence limit, the cue distribution is considered to have stabilized, and the gain from continued iteration is no longer significant, at which point the iteration loop terminates. The final defect detection information is the current detection result output in the last round at the end of the iteration, containing defect location markers and defect category markers after multiple rounds of adaptive optimization of the cue distribution. Its detection accuracy is significantly improved compared to the detection results using only the initial cue distribution.
[0139] In one embodiment, step S500 may specifically include the following steps S510 to S560: Step S510: Based on the defect location markers contained in the current detection results, extract the local image region corresponding to the defect region on the complete enhanced image, and perform Fourier transform on the local image region to obtain the amplitude spectrum representation and phase spectrum representation of the defect region.
[0140] The local image region corresponding to the defect area is a rectangular sub-image cropped from the complete enhanced image based on the bounding box coordinates of each defect location marked in the current detection result. The cropping includes the bounding box region and its outwardly extending narrow boundary region to avoid spectral leakage caused by boundary truncation. The same image block Fourier transform process as in step S100 is performed on each local image region: if the local image region size differs from the standard block size, it is first extended to the standard block size through edge mirroring; then, a two-dimensional discrete Fourier transform, quadrant centering shift, amplitude extraction, and phase extraction are performed to obtain the amplitude spectrum representation and phase spectrum representation of the defect region for each defect region. If the current detection result contains multiple defects, the amplitude spectrum representation and phase spectrum representation are extracted for each defect region separately.
[0141] Step S520: Compare the amplitude spectrum representation of the defective region with the amplitude spectrum distribution of the ideal defect-free region. Generate an amplitude indication error based on the frequency domain distribution of the difference. The amplitude indication error records the current amplitude indication distribution at the position where the defective frequency band is insufficiently suppressed and the background frequency band is excessively suppressed.
[0142] The ideal defect-free amplitude spectrum distribution is a reference model of the amplitude spectrum that a panel should have in a defect-free state, reflecting the frequency domain energy distribution of an ideal image without defects. By comparing the amplitude spectrum representation of the defective region with this ideal distribution frequency by frequency, and calculating the direction and amount of deviation between the actual amplitude value and the ideal amplitude range, it is possible to determine at which frequency positions the current amplitude indication distribution is insufficient in suppressing defects and at which frequency positions it is excessive in suppressing normal textures, thereby generating the amplitude indication error.
[0143] In one embodiment, step S520 may specifically include the following steps S521 to S526: Step S521: Establish a statistical model of the amplitude spectrum of an ideal defect-free panel image. The statistical model records the expected amplitude range of each frequency component under defect-free conditions.
[0144] The statistical model for the amplitude spectrum of an ideal defect-free panel image is established based on the standard amplitude spectrum set obtained in step S210. For each frequency point in the standard amplitude spectrum set, the amplitude value sequence of all defect-free images is extracted, and the mean and variance of the sequence are calculated. Using the mean as the center and an integer multiple of the standard deviation as the half-width, the expected amplitude range at that frequency point is constructed, ensuring that this expected range can cover the amplitude fluctuations of a normal panel at that frequency point with a preset confidence level. The expected amplitude range values for each frequency point are recorded as a range lookup table with the same size as the amplitude spectrum of the image patch. Each element in the table stores two values: an upper limit value and a lower limit value.
[0145] Step S522: Compare the amplitude of each frequency component in the amplitude spectrum representation of the defect region with the expected amplitude range of the corresponding frequency, and extract the positive deviation frequency component that exceeds the upper limit of the expected amplitude range and the negative deviation frequency component that is below the lower limit of the expected amplitude range. The positive deviation frequency component corresponds to the energy enhancement position caused by the appearance of the defect, and the negative deviation frequency component corresponds to the energy weakening position caused by excessive background suppression.
[0146] The comparison process involves traversing each frequency point represented by the amplitude spectrum of the defective region, extracting the actual amplitude at that frequency point, and comparing it with the upper and lower limits of the expected amplitude range for the same frequency point in the statistical model. If the actual amplitude is greater than the upper limit of the expected amplitude range, it indicates that there is additional energy enhancement introduced by the defect at that frequency point, and this frequency point is recorded as a positive deviation frequency component, with the excess amplitude difference recorded as a positive deviation. If the actual amplitude is less than the lower limit of the expected amplitude range, it indicates that the amplitude at that frequency point is weakened to below the normal level due to excessive suppression of the background texture by the current amplitude cue distribution, and this frequency point is recorded as a negative deviation frequency component, with the missing amplitude difference recorded as a negative deviation. Frequency points whose actual amplitude falls within the expected amplitude range do not generate deviation records.
[0147] Step S523: Based on the ratio of the amplitude excess of the positive deviation frequency component to the upper limit of the expected amplitude range of the corresponding frequency, generate a first amplitude adjustment indicator at the frequency position. The first amplitude adjustment indicator is used to guide the current amplitude indication distribution to increase the degree of suppression at the frequency.
[0148] For each positive deviation frequency component, the ratio of its excess amplitude difference to the upper limit of its expected amplitude range is calculated. This ratio reflects the relative strength of the energy enhancement caused by the defect relative to the normal upper limit. This ratio is mapped to a first amplitude adjustment indicator through a preset mapping function. The mapping function ensures that the larger the ratio, the larger the first amplitude adjustment indicator, and the increasing trend decreases to avoid overcorrection. A positive first amplitude adjustment indicator indicates that the current amplitude feedback at that frequency position needs to further reduce the response, i.e., increase the suppression depth.
[0149] Step S524: Based on the ratio of the insufficient amplitude of the negative deviation frequency component to the lower limit of the expected amplitude range of the corresponding frequency, generate a second amplitude adjustment indicator at the frequency position. The second amplitude adjustment indicator is used to guide the current amplitude indication to withdraw part of the original suppression degree at the frequency.
[0150] For each negative deviation frequency component, the ratio of its deficit amplitude difference to the lower limit of its expected amplitude range is calculated. This ratio reflects the relative degree of excessive background suppression. This ratio is mapped to a second amplitude adjustment indicator through a preset mapping function. The mapping function also ensures that the larger the ratio, the larger the indicator, with a decreasing growth trend. A negative second amplitude adjustment indicator indicates that the current amplitude at that frequency location needs to increase the response towards the reference flux, i.e., partially withdraw the suppression.
[0151] Step S525: Fill the first amplitude adjustment indicator and the second amplitude adjustment indicator at each frequency position into the frequency domain adjustment template of the same size as the amplitude indication distribution according to the frequency domain coordinates to obtain the amplitude adjustment distribution map.
[0152] The frequency domain adjustment template is an initial matrix of all zeros, with the same size and amplitude indication distribution as the template. Iterate through all positive deviation frequency components, filling the corresponding frequency point coordinates with the first amplitude adjustment indication value for each. Then iterate through all negative deviation frequency components, filling the corresponding frequency point coordinates with the second amplitude adjustment indication value for each. Frequency points not involved in the adjustment remain at zero. After filling, an amplitude adjustment distribution map is obtained. Positive value areas in the map represent locations where suppression needs to be deepened, negative value areas represent locations where suppression needs to be withdrawn, and zero value areas represent locations where no adjustment is needed.
[0153] Step S526: Perform frequency domain smoothing on the amplitude adjustment distribution map to eliminate abrupt changes in adjustment between adjacent frequency components, and obtain the amplitude indication error amount for superposition.
[0154] Frequency domain smoothing employs a two-dimensional Gaussian smoothing kernel to perform convolution filtering on the amplitude adjustment distribution map. The standard deviation of the Gaussian smoothing kernel is set relatively small to ensure that the smoothing range is limited to the local neighborhood, smoothing out abrupt peaks without disrupting the large-scale distribution pattern of the positive and negative adjustment regions. The smoothed amplitude adjustment distribution map is the amplitude indication error, which will be used to update the current amplitude indication distribution.
[0155] Step S530: Compare the phase spectrum representation of the defective region with the ideal defect-free phase distribution, generate a phase indication error based on the spatial frequency distribution of the phase deviation, and record the current phase indication distribution at the position where the defective phase adjustment is insufficient.
[0156] An ideal, defect-free phase distribution is a phase spectrum reference model that the panel should have in a defect-free state. The phase spectrum representation of the defective area is compared with this ideal distribution at frequency points to evaluate the phase deviation. The frequency components in the defective area whose phase deviation exceeds the normal tolerance range are obtained, and additional phase offset enhancement indicators are generated for these components. These are then integrated into a phase indication error quantity.
[0157] In one embodiment, step S530 may specifically include the following steps S531 to S536: Step S531: Call the standard phase distribution model of the defect-free panel. The standard phase distribution model of the defect-free panel records the expected phase interval of each frequency component in the normal panel.
[0158] The standard phase distribution model for a defect-free panel is constructed based on the standard phase spectrum set and the set of phase-stable frequency points generated in step S240. For different frequency points, the model records two types of information: for phase-stable frequency points, the model records the expected phase center value and the allowable fluctuation half-width under defect-free conditions, forming a narrow expected phase interval; for non-phase-stable frequency points, the model records a wider default allowable phase interval to adapt to the natural phase variation range under normal conditions.
[0159] Step S532: Determine the degree of deviation between the phase value of each frequency component represented by the phase spectrum of the defective region and the desired phase interval, obtain the phase deviation of each frequency component, and perform periodic adjustment on the phase deviation to make it fall into the main value interval.
[0160] The process of determining the degree of deviation is as follows: Calculate the cyclic difference between the actual phase value of the defective region and the center value of the desired phase interval. This difference is minimized by multiples of ±2π to obtain the cyclic phase difference. If this cyclic phase difference is within half the width of the positive and negative desired phase intervals, the frequency point is considered to have normal phase, and the phase deviation is recorded as zero. If the cyclic phase difference exceeds half the width of the desired phase interval, the excess difference is taken as the phase deviation; a positive value is recorded for deviations exceeding the positive half-width, and a negative value for deviations exceeding the negative half-width. Periodic adjustment is performed on the phase deviation. If the absolute value of the deviation is less than π, no operation is needed; if the absolute value exceeds π, it is adjusted by ±2π to fall within the range of negative π to positive π.
[0161] Step S533: Select frequency components whose phase deviation exceeds the preset tolerance boundary as phase abnormal frequency components. The phase abnormal frequency components are the positions where the defect causes phase changes.
[0162] The preset tolerance boundary is a positive threshold larger than half the width of the desired phase interval. It is used to filter out minor phase deviation noise, identifying only the more significant phase deviation frequency components as true phase anomalies caused by defects. All frequency points are iterated through, and frequency components whose absolute phase deviation value is greater than the preset tolerance boundary are recorded as phase anomaly frequency components, while retaining both the numerical value and sign of their phase deviation.
[0163] Step S534: Based on the difference between the phase deviation of each abnormal frequency component and the preset tolerance boundary, determine the additional phase offset depth required for that frequency component and generate a phase offset deepening indicator.
[0164] When calculating the phase offset deepening indicator, the net excess deviation is obtained by subtracting a preset tolerance boundary from the absolute value of the phase deviation. This net excess deviation is then multiplied by a preset gain coefficient to obtain the magnitude of the phase offset deepening indicator, which has the same sign as the phase deviation. The gain coefficient controls the step size of the phase offset correction in a single iteration. A positive phase offset deepening indicator indicates that the phase perturbation needs to be deepened in the positive direction based on the original offset, while a negative indicator indicates that the phase perturbation needs to be deepened in the opposite direction.
[0165] Step S535: Map the phase offset deepening indicator corresponding to each phase abnormal frequency component to an error construction plane of the same size as the phase indication distribution according to the frequency domain coordinates to obtain the original phase error distribution map.
[0166] The error construction plane is an initial all-zero matrix with the same size as the phase indication distribution. The phase shift intensification indicator for each phase-abnormal frequency component is filled into the corresponding frequency coordinate position of this matrix, while the frequency points where no phase abnormality occurs remain at zero, thus constructing the original phase error distribution map.
[0167] Step S536: Perform conformal filtering on the original phase error distribution map to filter out isolated error indicators while preserving the clear boundaries of the error concentration area, and obtain the phase indication error amount for superposition. The phase indication error amount is used to correct the offset of the corresponding frequency position in the current phase indication distribution.
[0168] Conformal filtering employs an edge-preserving filter, which smooths flat internal regions while maintaining clear edges in areas of significant error. The conformal filter uses a bilateral filtering algorithm, which, when calculating the filtered output at each location, considers both spatial proximity weights and similarity weights of error amplitudes at the same frequency point. Spatial proximity weights ensure smoothing occurs within a local window, while similarity weights prevent smoothing from crossing boundaries with significant differences in error amplitudes. The filter parameter settings allow isolated, small-area error indicators to be filtered out while large, contiguous error regions retain their clear outlines. The original phase error distribution map after conformal filtering represents the phase indication error.
[0169] Step S540: The amplitude indication error is superimposed on the current amplitude indication distribution to enhance and correct the suppression degree of the corresponding frequency point in the current amplitude indication distribution. At the same time, the phase indication error is superimposed on the current phase indication distribution to deepen the offset of the corresponding frequency point in the current phase indication distribution, thus obtaining the corrected amplitude indication distribution and the corrected phase indication distribution.
[0170] The amplitude indication error is additively superimposed with the current amplitude indication distribution position by position. After superposition, values exceeding the reasonable response range are truncated: values below a preset minimum response limit are truncated to that limit, and values above the baseline flux are truncated to the baseline flux. The truncated amplitude indication distribution is the corrected amplitude indication distribution, which has a stronger suppression depth in the defective frequency band compared to the uncorrected one, and recovers some flux in the over-suppressed background frequency band. The phase indication error is also additively superimposed with the current phase indication distribution position by position. After superposition, the upper limit of the phase offset amplitude is truncated to prevent excessive offset from causing image reconstruction distortion. The truncated result is the corrected phase indication distribution, which has a deeper offset at the phase aberration frequency components.
[0171] Step S550: Replace the current amplitude cue distribution and phase cue distribution with the corrected amplitude cue distribution and the corrected phase cue distribution. Based on the replaced cue distribution, perform all processing steps from applying the amplitude cue distribution to the amplitude spectrum representation to outputting the current detection result again to generate a new round of current detection results.
[0172] The system cache is updated with the corrected amplitude cue distribution values for the current image patch group position. Similarly, the phase cue distribution values are updated with the corrected phase cue distribution values. Using the updated cue distributions, all operations—including the invocation of amplitude and phase cue distributions, inverse Fourier transform reconstruction, image patch stitching, and inference by the pre-trained detection model—are re-executed for the panel image to be detected, starting from step S200, generating a new round of detection results. The new results will reflect the impact of the cue distribution correction on defect detection performance.
[0173] Step S560: Repeat the iterative process of extracting the amplitude spectrum representation and phase spectrum representation of the defect region until a new round of current detection results is generated. When the change in the cross-union ratio of the defect position marker between two consecutive iterations is lower than the preset convergence limit, it is determined that the termination condition for prompting adjustment is met, the iteration is stopped and the final defect detection information is output.
[0174] The iterative process begins at step S510 and ends at step S550, forming a complete adaptive correction loop for the cue distribution. After each iteration, the change in the cross-union ratio (CUNR) between the defect location markers in the current iteration and those in the previous iteration is calculated. The CUNR is calculated by pairing all defect bounding boxes from both iterations using the nearest neighbor principle. The area of the intersection of the two bounding boxes is divided by the area of their union to obtain the CUNR. The average of all paired CUNRs is used as the CUNR index for the current iteration, and the absolute value of the difference between the CUNR indices of two adjacent iterations is used as the CUNR change. A preset convergence limit is a small positive number. When the CUNR change falls below this limit, it indicates that the detection results between the two iterations are spatially consistent, the effect of the cue distribution adjustment has converged, and further iteration no longer produces significant location optimization. At this point, the iteration loop terminates. The final detection result of the last round at the end is output as the final defect detection information of the panel image to be detected. The information includes the final defect location mark and the final defect category mark. The two constitute the final output report of the panel defect detection system.
[0175] Figure 3 A hardware entity diagram of a computer system provided as an embodiment of the present invention, such as... Figure 3 As shown, the hardware entity of the computer system 1000 includes a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can run on the processor 1001, and the processor 1001 executes the program to implement the steps in the method of any of the above embodiments.
[0176] The memory 1002 stores computer programs that can run on the processor. The memory 1002 is configured to store instructions and applications that can be executed by the processor 1001. It can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) of the processor 1001 and various modules in the computer system 1000. It can be implemented by flash memory or random access memory (RAM).
[0177] When the processor 1001 executes the program, it implements the steps of the panel defect detection method based on visual cue learning described above. The processor 1001 typically controls the overall operation of the computer system 1000.
[0178] This invention provides a computer storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the panel defect detection method based on visual cue learning as described in any of the above embodiments.
[0179] It should be noted that the descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the storage medium and device embodiments of the present invention, please refer to the descriptions of the method embodiments of the present invention for understanding. The processor described above can be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that the electronic device implementing the above processor function can also be other types, and the embodiments of the present invention do not specifically limit it.
[0180] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0181] The above description is merely an embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A panel defect detection method based on visual cue learning, characterized in that, include: The panel image to be detected is divided into multiple image blocks. A Fourier transform is performed on each image block to obtain the amplitude spectrum representation and phase spectrum representation of the image block. The amplitude spectrum represents the periodic component distribution of the panel texture, and the phase spectrum represents the spatial phase relationship of the panel structure. A pre-constructed amplitude cue distribution and phase cue distribution are invoked. The amplitude cue distribution is applied to the amplitude spectrum representation to selectively suppress the spectral energy corresponding to the background texture, resulting in an enhanced amplitude spectrum representation. Simultaneously, the phase cue distribution is applied to the phase spectrum representation to adjust the phase distribution of the defective region, resulting in an enhanced phase spectrum representation. The amplitude cue distribution is a two-dimensional weight matrix with the same frequency-domain plane size as the image patch amplitude spectrum representation. Each element in the matrix records the amplitude adjustment indication at the corresponding frequency point, and the magnitude of the indication determines the degree to which the spectral energy at that frequency point is preserved or suppressed. The phase cue distribution is a two-dimensional offset matrix with the same frequency-domain plane size as the image patch phase spectrum representation. Each element in the matrix records the phase offset indication at the corresponding frequency point, and the magnitude of the offset determines the magnitude of the phase value adjustment at that frequency point. Perform an inverse Fourier transform on the enhanced amplitude spectrum representation and the enhanced phase spectrum representation of each image block to reconstruct the enhanced image block, and then stitch all the enhanced image blocks together according to the division and arrangement order to obtain the complete enhanced image; The pre-trained detection model with fixed response characteristics is input to the complete enhanced image. The pre-trained detection model performs defect localization and category prediction on the complete enhanced image and outputs the current detection result, which includes defect location markers and defect category markers. Based on the current detection result, a frequency domain prompt adjustment amount is constructed. The frequency domain prompt adjustment amount is an error feedback signal that performs directional correction on the amplitude prompt distribution and the phase prompt distribution according to the current detection result. The amplitude prompt distribution and the phase prompt distribution are corrected by the frequency domain prompt adjustment amount to obtain the corrected amplitude prompt distribution and the corrected phase prompt distribution. Based on the corrected amplitude prompt distribution and the corrected phase prompt distribution, the process of calling the pre-constructed amplitude prompt distribution and the phase prompt distribution to output the current detection result is re-executed until the prompt adjustment termination condition is met. The defect location mark and defect category mark of the last output are used as the final defect detection information.
2. The panel defect detection method based on visual cue learning according to claim 1, characterized in that, The process involves calling a pre-constructed amplitude and phase cue distributions, applying the amplitude cue distributions to the amplitude spectrum representation to selectively suppress the spectral energy corresponding to the background texture, resulting in an enhanced amplitude spectrum representation; and simultaneously applying the phase cue distributions to the phase spectrum representation to adjust the phase distribution of the defective regions, resulting in an enhanced phase spectrum representation. This includes: Multiple standard images of defect-free panels are acquired, and a Fourier transform is performed on each standard image of defect-free panels to obtain a standard amplitude spectrum set and a standard phase spectrum set. The standard amplitude spectrum set is subjected to frequency point aggregation processing to obtain the standard amplitude spectrum aggregation result. In the standard amplitude spectrum aggregation result, frequency points with energy concentration exceeding a preset limit are identified, and the identified frequency points are determined as the main frequency point set of the background texture. An amplitude suppression distribution is constructed based on the set of main frequency points of the background texture. The amplitude suppression distribution has a response amount less than the reference flux at each frequency point within the set of main frequency points of the background texture, and a response amount of the reference flux at frequency points outside the set of main frequency points of the background texture. The constructed amplitude suppression distribution is used as the amplitude cue distribution. Phase deviation statistics are performed on the standard phase spectrum set to determine the phase fluctuation range of each frequency point position between the defect-free panel standard images. The frequency points whose phase fluctuation range is less than the preset stability limit are determined as the phase stable frequency point set. A phase perturbation distribution is generated based on the phase stable frequency point set. The phase perturbation distribution carries a preset offset at the frequency points outside the phase stable frequency point set, and the generated phase perturbation distribution is used as the phase prompt distribution. The amplitude spectrum representation of each image block is combined with the amplitude cue distribution position by position. The amplitude at the set of main frequency points of the background texture is compressed based on the response amount less than the reference flux in the amplitude cue distribution. The amplitude at other frequency points is retained based on the response amount of the reference flux in the amplitude cue distribution, thus obtaining the enhanced amplitude spectrum representation. The phase spectrum representation of each image block is combined with the phase cue distribution position by position. The phase values at frequency points outside the phase stable frequency point set are adjusted based on a preset offset in the phase cue distribution, while the phase values within the phase stable frequency point set are retained, resulting in an enhanced phase spectrum representation.
3. The panel defect detection method based on visual cue learning according to claim 2, characterized in that, The step of constructing an amplitude suppression distribution based on the set of dominant frequency points of the background texture, wherein the amplitude suppression distribution has a response amount less than the reference flux at each frequency point within the set of dominant frequency points of the background texture, and has a response amount of the reference flux at frequency points outside the set of dominant frequency points of the background texture, and using the constructed amplitude suppression distribution as the amplitude cue distribution, includes: Each frequency point in the set of main frequency points of the background texture is mapped to a suppression control unit at the corresponding position in the amplitude suppression distribution, and an initial suppression degree is assigned to each suppression control unit; Perform a main frequency energy attenuation trend analysis on the standard amplitude spectrum aggregation result. Based on the attenuation trend of energy at each background texture main frequency point with frequency change, adjust the suppression degree of the corresponding suppression control unit so that the suppression degree is correlated with the attenuation trend in the same direction. The adjusted suppression control unit is smoothed in the neighborhood, and the suppression degree of each suppression control unit is harmonized with the suppression degree of the suppression control unit at its adjacent frequency point, so that the suppression degree transitions continuously in the frequency domain. A full-pass control unit is set at a non-dominant frequency position of the amplitude suppression distribution, and a reference flux value is assigned to the full-pass control unit so that the amplitude at the non-dominant frequency position is preserved. The entire suppression control unit and the all-pass control unit are combined into an overall amplitude suppression distribution. The overall amplitude suppression distribution has a continuously changing response in the frequency domain and the central region corresponds to the set of main frequency points of the background texture. This overall amplitude suppression distribution is determined as the amplitude prompt distribution. The amplitude indication distribution is subjected to frequency domain range constraint processing. While maintaining the degree of suppression of the main frequency point, the frequency domain span of the suppression transition band is reduced, so that the suppression transition region of the amplitude indication distribution is concentrated around the set of main frequency points of the background texture.
4. The panel defect detection method based on visual cue learning according to claim 2, characterized in that, The process involves performing phase deviation statistics on the standard phase spectrum set to determine the phase fluctuation range of each frequency point within the standard images of the defect-free panel. Frequency points with phase fluctuation ranges smaller than a preset stability limit are identified as a set of phase-stable frequency points. A phase perturbation distribution is generated based on this set of phase-stable frequency points. This phase perturbation distribution carries a preset offset at frequency points outside the set of phase-stable frequency points. The generated phase perturbation distribution is used as the phase indication distribution, including: Obtain the phase values of each defect-free panel standard image in the standard phase spectrum set at the same frequency point to obtain the phase value sequence corresponding to each frequency point; The discreteness of the phase value sequence corresponding to each frequency point is statistically analyzed to obtain the phase discreteness quantity. Frequency points with phase discreteness quantities less than a preset stability threshold are labeled as phase stable frequency points. All phase stable frequency points are collected to form a phase stable frequency point set. The spatial connected domains of the set of phase-stable frequency points in the frequency domain plane are determined. Extrapolation and expansion are performed on each spatial connected domain to obtain the boundary of the phase protection zone. The frequency points inside the boundary of the phase protection zone maintain the original phase relationship, while the frequency points outside the boundary are subjected to phase shift processing. In the phase disturbance distribution, a preset offset is filled into each frequency point outside the boundary of the phase protection zone. The amplitude of the preset offset is monotonically increasing with the shortest distance from the frequency point to the boundary of the phase protection zone. The frequency points inside the phase protection zone boundary are filled with zero offset in the phase disturbance distribution, so that the phase disturbance distribution presents a zero value region inside and a gradually increasing offset transition band outside. The filled-in phase perturbation distribution is used as the phase cue distribution, which is used to perform position-by-position combination processing with the phase spectrum representation of each image block when called.
5. The panel defect detection method based on visual cue learning according to claim 1, characterized in that, The enhanced amplitude spectrum representation and the enhanced phase spectrum representation of each image block are subjected to inverse Fourier transform to reconstruct the enhanced image block. All enhanced image blocks are then stitched together according to the partitioned arrangement order to obtain a complete enhanced image, including: The enhanced amplitude spectrum representation and enhanced phase spectrum representation of each image block are combined into a complex spectrum representation. The real part of the complex spectrum representation is determined by the correspondence between the cosine components of the enhanced amplitude spectrum representation and the enhanced phase spectrum representation, and the imaginary part is determined by the correspondence between the sine components of the enhanced amplitude spectrum representation and the enhanced phase spectrum representation. A two-dimensional discrete Fourier inverse transform is performed on the complex spectral representation to convert the frequency domain complex information to the spatial domain, thereby obtaining the initial reconstructed image block of the image block; The initial reconstructed image block is subjected to dynamic range stretching to map the pixel value distribution to a preset visible range, maintaining the relative contrast level between the defective area and the background, and generating an enhanced image block. The placement coordinates of each enhanced image block in the complete image are determined according to the horizontal and vertical arrangement numbers of the enhanced image blocks in the panel image to be detected. Transition fusion processing is performed at the stitching boundary between adjacent enhanced image blocks. Based on the pixel change gradient on both sides of the boundary between two adjacent enhanced image blocks, a gradient blending weight is generated to make the pixel values at the stitching seam transition smoothly. All enhanced image blocks that have undergone splicing boundary transition blending are combined into a complete enhanced image, the size of which is the same as the panel image to be detected.
6. The panel defect detection method based on visual cue learning according to claim 5, characterized in that, The step of performing a two-dimensional discrete Fourier inverse transform on the complex spectral representation to convert the frequency domain complex information to the spatial domain, thereby obtaining the initial reconstructed image patch of the image patch, includes: Perform a two-dimensional discrete Fourier inverse transform on the complex spectral representation to obtain the spatial domain real matrix of the image patch as the image to be reconstructed; The response differences of each frequency point in the amplitude indication distribution relative to the reference flux are obtained. The response differences are arranged along the frequency domain coordinates to form a frequency domain suppression depth record. An inverse Fourier transform is performed on the frequency domain suppression depth record to obtain a spatial domain suppression intensity distribution map. The pixel positions of the spatial domain suppression intensity distribution map correspond one-to-one with the pixel positions of the reconstruction map to be processed. In the spatial domain suppression intensity distribution map, continuous areas with suppression intensity below a preset intensity limit are defined as texture preservation zones, and continuous areas with suppression intensity not below a preset intensity limit are defined as defect highlighting zones. Based on the boundaries of the texture preservation partition and the defect highlighting partition, a binarized region segmentation template is constructed. The binarized region segmentation template takes a first marker value at the texture preservation partition position and a second marker value at the defect highlighting partition position. Using the binarized region segmentation template, the pixel set located in the defect highlighting zone of the image to be reconstructed is contrast-expanded to increase the gray-level span of the pixel set. At the same time, the gray-level distribution of the pixel set located in the texture preservation zone is preserved so that the gray-level statistical characteristics of the pixel set are consistent with those before processing. The imperfection highlighting pixels after contrast expansion and the texture preservation pixels after grayscale distribution preservation are merged according to their spatial location to obtain the initial reconstructed image block of the image block.
7. The panel defect detection method based on visual cue learning according to claim 5, characterized in that, The transition fusion process at the stitching boundary between adjacent enhanced image blocks, based on the pixel change gradient on both sides of the boundary between two adjacent enhanced image blocks, generates a gradient blending weight to ensure a smooth transition of pixel values at the stitching seam, including: Extract the inner pixel strips of the first enhanced image block and the second enhanced image block adjacent to the current stitching boundary. The width of the two pixel strips corresponds to the set half-width of the fusion transition region. The brightness gradient distribution of the pixel strips of the first enhanced image block along the direction perpendicular to the boundary and the brightness gradient distribution of the pixel strips of the second enhanced image block along the direction perpendicular to the boundary are determined respectively. The brightness gradient distributions on both sides are normalized and aligned to make the gradient change rates of the two enhanced image blocks at the boundary consistent, thus generating a brightness gradient matching curve. A gradient weight function is constructed based on the brightness gradient matching curve. The gradient weight function assigns the weight of the first enhanced image block as the dominant weight at the beginning of the fusion transition region, assigns the weight of the second enhanced image block as the dominant weight at the end, and transitions in the middle according to a gradient rule. Based on the gradient weighting function, each pixel in the fusion transition region of the first enhanced image block and the second enhanced image block is mixed to generate a sequence of fusion boundary pixel values. The original pixels at the corresponding stitching boundary are replaced by the fusion boundary pixel value sequence, so that the first enhanced image block and the second enhanced image block form a continuous image region along the stitching boundary, thus forming an enhanced image block combination after eliminating the block effect.
8. The panel defect detection method based on visual cue learning according to claim 1, characterized in that, The pre-trained detection model, which fixes the input response characteristics of the complete enhanced image, performs defect localization and category prediction on the complete enhanced image and outputs the current detection result. The current detection result includes defect location markers and defect category markers, including: The complete enhanced image is fed into the deep convolutional processing stream of the pre-trained detection model. The image content is extracted step by step through multiple convolutional processing layers configured in series, generating a multi-level abstract representation map set. The receptive field coverage of each abstract representation map increases as the layer deepens. Extract the deepest level of the global semantic abstract representation graph from the multi-level abstract representation graph set, perform region activation mapping on the global semantic abstract representation graph, and locate the suspected defective activation regions whose response intensity exceeds the preset response threshold. Select at least two shallow-level detail abstraction maps from the multi-level abstraction map set, align the shallow-level detail abstraction maps with the suspected defect activation areas in spatial position, and crop out the local detail representations corresponding to each suspected defect activation area; For each suspected defective activation region, perform bounding box offset regression on the local detail representation to generate a corrected, precise location bounding box; Perform category probability mapping on the local detail representation within each precise location bounding box, mapping the local detail representation to a probability distribution on a preset set of defect categories, and selecting the defect category corresponding to the highest probability in the probability distribution as the defect category label for the bounding box; The coordinates of all precise location bounding boxes and their corresponding defect category labels are aggregated into the current detection result. The defect location labels in the current detection result are composed of the coordinate sequence of the precise location bounding boxes, and the defect category labels are composed of the defect category labels corresponding to each bounding box.
9. The panel defect detection method based on visual cue learning according to claim 8, characterized in that, The process of feeding the complete enhanced image into the deep convolutional processing stream of the pre-trained detection model, and refining the image content step by step through multiple sequentially configured convolutional processing layers to generate a multi-level abstract representation atlas, includes: Obtain the pre-fixed convolutional processing layer arrangement order in the pre-trained detection model. The convolutional processing layer includes alternately set filter expansion layers and compression aggregation layers. The increase of the filter expansion layer indicates the number of channels, and the reduction of the compression aggregation layer indicates the plane size. The complete enhanced image is input into the first convolutional processing layer, where filtering expansion and nonlinear activation are performed to obtain the initial level abstract representation map. The current level abstract representation is passed to the next convolutional processing layer, where compression and aggregation are performed to reduce the plane size, and further filtering and expansion are performed to increase the number of channels, resulting in the next level abstract representation. This passing process is repeated until all convolutional processing layers are traversed. After each compression aggregation level, the planar size of the abstract representation map output by the compression aggregation level becomes the preset compression ratio of the planar size of the input abstract representation map. At the same time, the image structure granularity captured by each channel of the output abstract representation map is coarser than that of the input. Record the intermediate abstract representation graphs output by each convolutional processing layer to obtain a sequence of abstract representation graphs sorted by processing depth. In the sequence of abstract representation graphs, shallow abstract representation graphs retain fine-grained geometric details, while deep abstract representation graphs capture a wide range of semantic layouts. The sequence of abstract representation graphs is used as a multi-level abstract representation graph set, and the spatial resolution and channel dimension of each abstract representation graph in the multi-level abstract representation graph set are distributed in a pyramid shape from shallow to deep.
10. A computer system comprising a memory and a processor, the memory storing a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Panel surface defect detection method, system and equipment and storage medium
CN118134895A
Shallow layer defect detection method and device based on phase residual error and storage medium
CN120471917A