A workpiece defect detection method and system based on machine vision
By using lightweight convolutional structures and attention modulation mechanisms, combined with depthwise separable convolution and feature clustering, the problem of identifying minute defects and providing early warnings in machine vision inspection is solved, enabling efficient and accurate detection of workpiece surfaces.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUXI INSTITUTE OF TECHNOLOGY
- Filing Date
- 2026-03-02
- Publication Date
- 2026-06-05
AI Technical Summary
Existing machine vision inspection solutions are unable to effectively identify minute defects on the surface of workpieces, especially defects of different shapes, and lack unified quantitative standards and early warning mechanisms, leading to missed detections and missed opportunities for quality intervention.
We employ a lightweight convolutional structure and attention modulation mechanism. We use depthwise separable convolution to extract features at multiple scales, combine attention enhancement processing to construct salient regions, perform candidate box screening and boundary regression, combine feature clustering and weak response analysis to achieve defect classification and early warning, and construct discrimination rules to output structured detection reports.
It enables effective perception and precise location of weak defect signals, accurately identifies defects of different morphologies, and promptly detects defects that have not yet fully developed through an early warning mechanism, thereby improving the accuracy and timeliness of detection.
Smart Images

Figure CN122156122A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial vision inspection technology, and in particular to a method and system for detecting workpiece defects based on machine vision. Background Technology
[0002] Surface quality inspection of workpieces is a crucial aspect of quality control in manufacturing, involving the identification and grading of appearance defects in various products such as metal castings, plastic injection molded parts, electronic components, and precision mechanical parts. With the advancement of industrial automation, machine vision-based automated inspection solutions are gradually replacing traditional manual visual inspection, becoming the mainstream technology for production line quality control.
[0003] However, existing machine vision inspection solutions still face several technical bottlenecks. Differences in the material texture and reflective properties of workpiece surfaces result in low contrast between some minute defects and the background. Conventional feature extraction methods can easily drown out weak defect signals in background noise, leading to missed detections. Furthermore, different defect morphologies, such as point pores, linear scratches, and surface corrosion, require different sizes of inspection frames to match, making it difficult for fixed templates to accommodate diverse defect contours. Simultaneously, there is a lack of unified quantitative standards for determining defect types and severity, and an effective early warning mechanism is lacking for early-stage defects that are still in their nascent stages. Often, defects are only identified after they have fully developed, missing the optimal opportunity for quality intervention. Summary of the Invention
[0004] This invention discloses a machine vision-based method and system for workpiece defect detection. It aims to achieve effective perception of weak defect signals through lightweight convolutional structures and attention modulation mechanisms, accurately locate defects of different shapes using morphological adaptive boundary regression, achieve joint evaluation of defect classification and early warning based on feature clustering and weak response analysis, and finally output a structured inspection report by constructing discrimination rules through feature sensitivity weighting, thus providing complete visual inspection capabilities for automated quality control of industrial production lines.
[0005] The first aspect of this invention proposes a machine vision-based method for detecting workpiece defects, comprising the following steps: Collect image sequence data of the workpiece surface, and perform size normalization and data augmentation on the image sequence data to form a preprocessed image set; For the preprocessed image set, multi-scale feature extraction is performed using depthwise separable convolution to generate feature maps. Attention enhancement processing is performed on the feature maps to construct salient regions. Based on the salient regions, candidate boxes are selected to generate defect response parameters. Spatial pyramid pooling is performed on the feature map to determine fusion features. The fusion features are used to perform bounding box regression to locate defect segments. Based on the defect segment features, the defect response parameters are fused to form a feature matrix. Pattern clustering is performed on the feature matrix to determine the category distribution. Weak response regions are extracted from the feature matrix to generate early warning markers. Defect level parameters are generated based on the category distribution and the early warning markers. Severity gradient features are extracted from the defect level parameters to construct a defect feature map. Feature sensitivity weights are extracted by performing feature optimization on the defect response parameters and the feature matrix through multi-scale loss constraints. A weighted discrimination rule is constructed based on the feature sensitivity weights. The defect feature map is then integrated through the weighted discrimination rule to output the visual detection result.
[0006] A second aspect of this invention provides a machine vision-based workpiece defect detection system, comprising: The image acquisition module is used to acquire image sequence data of the workpiece surface, and to perform size normalization and data enhancement processing on the image sequence data to form a preprocessed image set; The feature extraction module is used to perform multi-scale feature extraction on the preprocessed image set through depthwise separable convolution to generate feature maps, perform attention enhancement processing on the feature maps to construct salient regions, and perform candidate box filtering based on the salient regions to generate defect response parameters. The segment localization module is used to perform spatial pyramid pooling on the feature map to determine fusion features, use the fusion features to perform bounding box regression to locate defect segments, and fuse the defect response parameters based on the defect segment features to form a feature matrix. The map construction module is used to perform pattern clustering on the feature matrix to determine the category distribution, extract weak response regions from the feature matrix to generate early warning markers, generate defect level parameters based on the category distribution and the early warning markers, and extract severity gradient features from the defect level parameters to construct a defect feature map. The detection output module is used to perform feature optimization on the defect response parameters and the feature matrix through multi-scale loss constraints to extract feature sensitivity weights, construct weighted discrimination rules based on the feature sensitivity weights, and integrate the defect feature map through the weighted discrimination rules to output visual detection results.
[0007] The beneficial effects of this invention are reflected in the following points: First, by employing a depthwise separable convolutional structure to decouple spatial filtering and channel mixing, computational complexity is reduced. Channel compression factors are used to recalibrate the importance of each channel, forming lightweight features. Multi-scale cascaded fusion is combined to capture details of minute defects and perceive the overall impact of large-area defects. Furthermore, dual-path modulation of channel attention and spatial attention compensates and enhances weak response regions, ensuring that low-contrast defect signals such as fine pores and fatigue crack initiation textures in aluminum alloy die castings are preserved rather than being drowned out by background noise. Second, spatial pyramid pooling is used to fuse multi-level contextual information from global to local levels. Multi-specification anchor frame templates are used to match defects of different morphologies, such as point pores, linear scratches, and planar corrosion. Aspect ratio constraint regression is combined to finely adjust the bounding box shape, ensuring the positioning result closely matches the actual defect contour. Non-maximum suppression eliminates redundant detection boxes for the same target, enabling accurate spatial positioning of defects of varying morphologies, such as long scratches on steel plates and pitting on gears. Finally, feature clustering analysis is used to group defect samples with similar response patterns into the same cluster to achieve automatic type classification. Regions with low response intensity but consistent with defect evolution characteristics are identified from the feature matrix and early warning markers are generated after time-series stability verification. This enables early warning of nascent defects that have not yet fully developed, such as early pitting of bearing raceways and early chipping of tool surfaces. At the same time, the sensitivity weights of each feature dimension are extracted through multi-scale loss constraints with category balance weighting, so that rare defect types receive the same attention as common defects in the discrimination process. Attached Figure Description
[0008] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.
[0009] Figure 1 This is a schematic flowchart of a machine vision-based workpiece defect detection method according to the present invention.
[0010] Figure 2 This is a schematic diagram of the layout of the workpiece surface image acquisition system of the present invention.
[0011] Figure 3 This is a structural block diagram of a workpiece defect detection system based on machine vision according to the present invention.
[0012] Among them: 1-Industrial camera; 2-Lens optical axis; 3-Ring LED array; 4-Low-angle side light source; 5-Front diffuse light source; 6-Workpiece; 7-Inspection station; 8-Trigger signal generator. Detailed Implementation
[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0014] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0015] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0016] The technical solutions of the embodiments of this application will be described below.
[0017] like Figure 1 As shown, this embodiment of the invention provides a workpiece defect detection method based on machine vision, including the following steps S110-S150: Step S110: Collect image sequence data of the workpiece surface, and perform size normalization and data augmentation on the image sequence data to form a preprocessed image set.
[0018] Specifically, image sequence data of the workpiece surface is acquired. For example... Figure 2As shown, industrial camera 1 is mounted above inspection station 7. The optical axis 2 of the lens is perpendicular to the surface of workpiece 6 to reduce imaging geometric distortion. The resolution of industrial camera 1 is set to 2048×2048 pixels. At this resolution, a single pixel corresponds to a physical size of approximately 0.05mm on the surface of workpiece 6, and can distinguish minute defects with a diameter of 0.1mm or larger. Image sequence data is synchronously acquired and generated by trigger signal generator 8 when workpiece 6 passes through inspection station 7. Multiple frames of images are acquired for each workpiece 6 according to its surface area and inspection accuracy requirements. A typical configuration is to acquire 4 to 8 frames of images covering different areas or different lighting angles for a single workpiece 6. The light source employs a ring-shaped LED array 3, which includes a low-angle side light source 4 and a forward diffused light source 5. By adjusting the brightness ratio of the low-angle side light source 4 and the forward diffused light source 5, different types of defects such as scratches, pits, and cracks on the surface of the workpiece 6 can be highlighted. When the brightness ratio of the low-angle side light source 4 is set to 60% to 80%, the shadow contrast of scratches is enhanced. When the brightness ratio of the forward diffused light source 5 is set to 70% to 90%, the display effect of color difference and stain defects is improved. The image sequence data is stored in a lossless compression format to retain the original pixel information. The data size of a single frame image is approximately 8 to 12 MB, and the total amount of image sequence data generated per hour in a continuous production environment can reach tens of GB.
[0019] Image sequence data undergoes size normalization and data augmentation to form a preprocessed image set. Size normalization uniformly scales the width and height of each frame in the image sequence data. Before scaling, the original dimensions of each frame are read and the scaling factor relative to the target size is calculated. The target size for size normalization is set to 512×512 pixels or 640×640 pixels. The scaling process uses a bicubic interpolation algorithm to maintain edge sharpness and detail texture. Images in the image sequence data whose aspect ratio does not match the target size are processed using either center cropping or edge filling strategies. The cropping strategy is suitable for scenarios where the workpiece is centered and the background area is large, while the filling strategy is suitable for scenarios where the workpiece occupies the main part of the image and edge information cannot be discarded. Data augmentation expands the size-normalized images with various enhancement operations, including four basic transformations: random rotation, horizontal flipping, random brightness transformation, and contrast adjustment. The rotation angle range is set to ±15 degrees to simulate random deviations in the workpiece loading posture, and the random brightness transformation range is set to ±20% to simulate illuminance fluctuations caused by light source aging or ambient light interference. The preprocessed image set is generated after data augmentation. Each original image frame is augmented and expanded into 3 to 5 transformed images, increasing the size of the preprocessed image set by 3 to 5 times compared to the image sequence data. Each image in the preprocessed image set is uniformly converted into a floating-point tensor and its pixel values are normalized. Each image contains three color channels: red, green, and blue. The normalization formula is I_norm = I_raw / 255, where I_raw is the original pixel value ranging from 0 to 255, and I_norm is the normalized pixel value ranging from 0 to 1. The normalized tensor data is directly used as input to the feature extraction network.
[0020] Step S120: For the preprocessed image set, multi-scale feature extraction is performed through depthwise separable convolution to generate feature maps. Attention enhancement processing is performed on the feature maps to construct salient regions. Candidate boxes are selected based on the salient regions to generate defect response parameters.
[0021] In some embodiments, the step of generating a feature map by multi-scale feature extraction using depthwise separable convolution on the preprocessed image set includes: performing channel-wise convolution on the preprocessed image set to generate channel feature responses; extracting cross-channel correlation information from the channel feature responses to generate channel compression factors; performing channel recalibration on the channel feature responses using the channel compression factors to form lightweight features; and performing multi-scale cascade fusion based on the lightweight features to generate a feature map.
[0022] Channel-wise convolution is performed on the preprocessed image set to generate channel feature responses. Each input channel of each image tensor in the preprocessed image set is configured with an independent 3×3 spatial convolution kernel. The convolution kernel independently slides within each channel to calculate the corresponding feature response map. The sliding stride is set to 1 pixel to preserve maximum spatial resolution, and zero-padding is used at the convolution boundaries to ensure that the output feature map has the same spatial size as the input image. The image tensors in the preprocessed image set have three input channels corresponding to the RGB primary colors. The R, G, and B channels each generate a feature response map, and the three response maps are stacked in the channel dimension to form a channel feature response tensor with the same number of input channels. Scratches on the surface of metal workpieces break through the surface coating or oxide layer, exposing the substrate material. These defects exhibit strong brightness jumps in the channel feature responses, with response values jumping from 0.2 to 0.3 in the background area to 0.7 to 0.9 in the scratch area. The response amplitude distribution of different color channels differs, and this differential response pattern between channels carries the color attribute information of the defect. After convolutional downsampling, the spatial resolution of the channel feature responses is reduced to 1 / 2 or 1 / 4 of the input image size. When the input image size is 512×512, the channel feature response size is 256×256 or 128×128. This reduced resolution expands the receptive field and reduces the amount of feature data. The channel feature responses retain the independent spatial filtering results of each channel, and the complementary features captured by different channels are integrated in the subsequent channel blending stage.
[0023] Channel compression factors are generated by extracting cross-channel correlation information from the channel feature response. Global average pooling compresses the two-dimensional feature map of each channel in the channel feature response into a scalar value, which represents the overall activation intensity of that channel. The compression process calculates the arithmetic mean of the values at all pixel positions in the feature map. After global average pooling, the 256×256 feature map is compressed into a single value, achieving a compression ratio of 65536:1. The channel compression factor is obtained by performing a fully connected transformation on the compressed channel activation vector. The fully connected transformation consists of two layers. The first layer compresses the number of channels C to C / 16 to reduce computation. The two layers introduce non-linear mapping capability through the ReLU activation function. The second layer restores the dimension to C to generate the importance coefficients of each channel. The fully connected transformation performs a non-linear mapping on the channel activation vectors to learn the correlation weights between channels. When there are C channels in the channel feature response, the channel compression factor is a C-dimensional vector, where the i-th element represents the importance coefficient of the i-th channel relative to the other channels. The numerical range of the channel compression factor is constrained to between 0 and 1 by the Sigmoid activation function. For defects primarily characterized by brightness variations, such as cracks and scratches, the importance coefficient of the brightness channel is close to 1 while that of the chroma channel is close to 0, because these defects mainly manifest as local dark or bright lines rather than color changes. For defects primarily characterized by color variations, such as color difference and oxidation discoloration, the importance coefficient of the chroma channel is close to 0.8 to 0.9 while that of the brightness channel decreases to 0.3 to 0.4, because these defects mainly manifest as color shifts rather than brightness variations. The channel compression factor is stored in a compact vector form. When C=64, the channel compression factor only occupies 64 floating-point numbers of storage space. The storage overhead is negligible compared to the complete channel feature response tensor, but it carries crucial prior information about the channel importance.
[0024] Lightweight features are formed by channel recalibration of the channel feature responses using a channel compression factor. Each element of the channel compression factor is used as a scaling weight and modulated element-wise with the corresponding channel in the channel feature response. All pixel values of the i-th channel are multiplied by the i-th element of the channel compression factor. Channels with high importance coefficients retain or enhance their feature amplitudes, while those with low importance coefficients are suppressed and approach zero. Lightweight features are formed after channel recalibration. The tensor shape of the lightweight features is completely consistent with the channel feature responses, but the feature amplitude distribution of each channel has changed significantly after modulation by the channel compression factor. When detecting oil stain defects on the surface of metal workpieces, channels reflecting surface gloss in the channel feature response are amplified to highlight the abnormal gloss caused by oil stains, with a magnification factor typically between 1.5 and 2.0. Meanwhile, channels reflecting the material's inherent color are suppressed to filter normal reflections, with a suppression factor typically between 0.1 and 0.3. When detecting porosity defects in castings, the edge response channel has a magnification factor of 2.0 to 2.5 to highlight the contour features of the pores, while the texture response channel has a suppression factor below 0.2 to filter out normal machining textures on the workpiece surface. Compared to unrecalibrated channel feature responses, lightweight features increase the inter-class distance between defect samples and normal samples in the feature space, thus improving the classification and discrimination ability of the detection network. The calculation process of lightweight features involves only element-wise multiplication operations, resulting in extremely low computational overhead and the ability to be performed in parallel with convolution operations.
[0025] Feature maps are generated through multi-scale cascaded fusion of lightweight features. The lightweight features are downsampled multiple times to generate feature pyramids with different spatial resolutions. Each pyramid contains four scale levels, with the resolution halved and the number of channels doubling sequentially. The bottom level has a resolution of 128×128 and 64 channels; the second level has a resolution of 64×64 and 128 channels; the third level has a resolution of 32×32 and 256 channels; and the top level has a resolution of 16×16 and 512 channels. Lightweight features have the largest receptive field coverage at the top of the pyramid, the smallest resolution level. The receptive field of a single pixel at the top covers the 32×32 pixel area of the original image, enabling the perception of large-scale defects spanning hundreds of pixels, such as the overall deformation of a metal workpiece. At the bottom of the pyramid, the largest resolution level, the finest spatial details are preserved, enabling the location of tiny defects occupying only a few pixels, such as microscopic pores on the surface of a casting. Cascaded fusion employs a top-down approach, progressively upsampling high-level semantic features and adding them to low-level detail features. Upsampling uses bilinear interpolation to double the feature map resolution. During fusion, 1×1 convolutions are used to align the number of channels at different levels to ensure element-wise dimensionality matching. The feature map is generated after cascaded fusion, with its spatial resolution set to the bottom layer of the pyramid to preserve maximum spatial detail. The channel dimension integrates semantic information from each level to form a multi-dimensional fused feature vector. The feature vector at each spatial location in the feature map simultaneously encodes both local texture details and global contextual semantics at that location.
[0026] In some embodiments, the step of performing attention enhancement processing on the feature map to construct a salient region includes: performing dual-path separation calculation of channel response and spatial response on the feature map to generate an initial attention map; extracting low-response channels from the initial attention map to generate suppression compensation weights; performing weak feature enhancement on the initial attention map using the suppression compensation weights to form a corrected attention map; and filtering high-response regions based on the corrected attention map to form a salient region.
[0027] The initial attention map is generated by performing dual-path calculations of channel response and spatial response on the feature map. Global average pooling along the spatial dimension of the feature map compresses the two-dimensional feature map of each channel into a single scalar. This scalar is then passed through two fully connected layers to output a channel attention vector of the same length as the number of channels. The intermediate dimension of the two fully connected layers is 1 / 8 of the number of channels. Spatial response calculation involves max pooling and average pooling along the channel dimension of the feature map, compressing the multi-channel features into two single-channel feature maps. These two feature maps are concatenated along the channel dimension and then modeled for spatial relationships using a 7×7 convolutional kernel. A shared convolutional layer learns the dependencies between spatial locations, outputting a spatial attention map with the same spatial resolution. The initial attention map is generated by fusing the channel attention vector and the spatial attention map through an outer product. This outer product operation fuses the C-dimensional channel vector with an H×W spatial matrix into a C×H×W three-dimensional tensor. The fused initial attention map has the same three-dimensional tensor shape as the feature map, with each element representing the comprehensive importance weight of the corresponding position in the corresponding channel. The attention-modulated features can be obtained by multiplying the feature map element by element with the initial attention map. However, directly using the initial attention map may lead to excessive suppression of effective features in weak response areas. The fine pores on the surface of the casting have weak response due to their reflective properties being close to those of the substrate. Direct suppression will lead to the missed detection of such defects.
[0028] Low-response channels are extracted from the initial attention map to generate suppression compensation weights. Channels and locations in the initial attention map with response intensities below a certain proportion of the mean are identified and marked. These low-response elements are given additional enhancement weights to prevent them from being completely suppressed. The identification threshold is set to 0.3 times the global mean. Some channels and locations in the initial attention map have attention weights close to zero. These low-response regions may correspond to background areas or early defect areas containing weak defect signals but misjudged as background due to their weak response. Early cracks on the surface of metal workpieces only exhibit weak surface texture anomalies in their initial stage, with response intensities far lower than those of mature cracks that have already expanded. The statistical distribution of response intensities in the initial attention map typically exhibits a long-tail characteristic. A small number of high-response elements, corresponding to significant defect areas, occupy the head of the distribution, while a large number of low-response elements, corresponding to background and weak defect areas, occupy the tail. The compensation threshold is set to the 25th percentile of the distribution to cover potential weak defects in the tail. The suppression compensation weight is set between 1 and 3. A weight of 1 indicates that no compensation is needed to maintain the original attention, while a weight greater than 1 indicates that enhancement is needed to prevent over-suppression. The weight value is negatively correlated with the response intensity; the lower the response intensity, the greater the compensation weight. Minor defects on the surface of the coated part receive a larger compensation weight due to their slight color difference to avoid missed detection. The suppression compensation weight has the same tensor shape as the initial attention map, and the two can be directly operated on element-wise to achieve compensation enhancement.
[0029] For example, the step of weakly enhancing the initial attention map to form a corrected attention map using the suppression compensation weight includes: performing regional difference analysis on the initial attention map based on the suppression compensation weight to generate an enhancement priority map; extracting weak response region boundaries from the enhancement priority map to generate boundary enhancement factors; performing adaptive contrast stretching on the initial attention map using the boundary enhancement factors to form an enhanced attention map; and performing response equalization processing on the enhanced attention map to form a corrected attention map.
[0030] An enhancement priority map is generated by performing regional difference analysis on the initial attention map based on suppression compensation weights. The difference between each low-response location in the initial attention map and the maximum response value in its neighborhood is calculated one by one. The neighborhood is set to a 5×5 pixel area centered on the current location. A larger difference indicates a more significant response gap between the location and its surrounding environment, and the corresponding location receives a higher priority value in the enhancement priority map. Suppression compensation weights identify the low-response regions in the initial attention map that need enhancement. Low-response locations located at the edge of high-response regions are more likely to contain defect boundary information and therefore receive higher priority values in the enhancement priority map. Isolated low-response locations far from any high-response regions are likely pure background and therefore receive lower priority values. Taking painted part inspection as an example, defect edge regions appear as low-response bands distributed along the defect direction in the initial attention map. These low-response regions are adjacent to the high-response regions of the defect body and therefore have high enhancement priority; while normal regions far from the defect, although also low-response, are unrelated to the defect and therefore have low enhancement priority. Locations with a suppression compensation weight greater than 1 receive a non-zero priority value in the enhancement priority map. The value is positively correlated with the difference in response in the neighborhood of that location. Low-response locations near the defect edge receive high priority because they form a significant difference with the adjacent high-response areas.
[0031] Boundary enhancement factors are generated by extracting the boundaries of weak response regions from the enhancement priority map. The gradients of priority values in the horizontal and vertical directions in the enhancement priority map are calculated using the Sobel operator. The horizontal Sobel kernel is [[-1,0,1],[-2,0,2],[-1,0,1]], and the vertical Sobel kernel is [[-1,-2,-1],[0,0,0],[1,2,1]]. Locations with larger gradient magnitudes are marked as boundary positions. Higher priority locations in the enhancement priority map are concentrated at the boundaries between high and low response regions. These locations correspond to the transition boundaries between the defect target and the background. Feature enhancement of these boundary regions can accurately locate the defect contour. The boundaries of planar defects on the surface of painted parts are usually blurred and gradual; if the boundary features are not sufficiently enhanced, the localization box will show significant deviations. Regions in the enhancement priority map where priority values change sharply from high to low generate strong gradient responses in edge detection. Locations with gradient magnitudes exceeding twice the global gradient mean are marked as boundary points and assigned larger boundary enhancement factor values. The boundary enhancement factor is set to a value between 1 and 2. A larger factor value at the boundary indicates a need for stronger enhancement, while a factor value of 1 at non-boundary locations indicates that standard enhancement is maintained. The boundary enhancement factor and the enhancement priority map have the same spatial resolution, and they can be used element-wise to achieve a differentiated enhancement strategy prioritizing boundaries. Although low-response locations in the non-boundary regions of the enhancement priority map also have some enhancement needs, their boundary enhancement factor is close to 1, resulting in a relatively moderate enhancement amplitude to avoid excessive amplification of background noise.
[0032] An enhanced attention map is generated by adaptive contrast stretching of the initial attention map using a boundary enhancement factor. The adaptive contrast stretching calculates the enhanced attention map by substituting the response values of the initial attention map into an S-curve stretching function, f(x) = 1 / (1 + exp(-k × (xm))), where x is the original response value at the stretching location in the initial attention map (range 0 to 1), k is a kurtosis parameter controlled by the boundary enhancement factor, m is the center point parameter (the median of the initial attention map), and f(x) is the enhanced response value at that location after stretching (range 0 to 1). The adaptive contrast stretching applies different intensities of stretching to different locations in the initial attention map. The stretching intensity is controlled by the boundary enhancement factor; a larger stretching intensity is applied to boundary locations to sharpen the defect contour, while a smaller stretching intensity is applied to non-boundary locations to maintain a smooth response. After strong stretching, the response values at the defect boundary locations in the initial attention map polarize, pushing the transitional responses that were originally in the middle grayscale to either high or low response ends, thus making the boundaries of minor scratches on the metal workpiece surface clearer and sharper. Compared to the initial attention map, the enhanced attention map exhibits a stronger bimodal characteristic in the overall response distribution, with the high response peak corresponding to the defect area, the low response peak corresponding to the background area, and the response proportion in the intermediate transition area significantly reduced.
[0033] A corrected attention map is generated by performing response equalization on the enhanced attention map. The enhanced attention map is then piecewise linearly mapped, with the following rules: response values below the lower threshold are truncated to the lower limit; response values above the upper threshold are truncated to the upper limit; and response values in the middle range are linearly scaled to the standard 0-1 range. After contrast stretching, the enhanced attention map may experience local response saturation, with some high-response regions exceeding a reasonable range, leading to loss of feature information. In severely corroded areas of metal workpieces, stretching may cause response values to overflow the effective range, smoothing out the gradient differences in corrosion severity. The corrected attention map is generated after response equalization. Its numerical range is strictly constrained to 0-1, exhibiting a clear bimodal distribution, with the valley between the peaks corresponding to the optimal binarization segmentation threshold. Regions saturated due to overstretching in the enhanced attention map are restored to reasonable values in the corrected attention map. Internal details of high-response regions are preserved rather than truncated to a uniform maximum value, which is valuable for distinguishing pit defects of different depths. As the final output of the attention enhancement process, the quality of the corrected attention map directly determines the accuracy of salient region extraction.
[0034] High-response regions are selected based on the corrected attention map to form salient regions. The global mean of the corrected attention map is multiplied by 1.5 to serve as an adaptive segmentation threshold. Typically, the threshold falls within the range of 0.45 to 0.55. Locations with weight values higher than the threshold are marked as foreground, i.e., potential defect regions, while locations with weight values lower than the threshold are marked as background. The weight value at each location in the corrected attention map reflects the confidence level that the location contains defect-related features. Locations with high weight values are likely to correspond to defect targets or defect edges, while locations with low weight values are likely to correspond to normal background textures. Foreground regions are merged into several independent salient regions after morphological connected component analysis. Porosity and crack defects on the surface of the casting form independent connected components within the salient regions. The geometric attributes of each connected component, such as area, aspect ratio, and roundness, are calculated and recorded simultaneously. These attributes are used to distinguish different types of defect patterns. In the corrected attention map, low-response regions enhanced by weak features are not included in the salient region if their weight after compensation is still below the threshold, indicating that the region is indeed background rather than a weak defect. If the weight after compensation exceeds the threshold, they are included in the salient region, indicating that the region contains a weak defect signal that was almost missed. The geometric attributes of the salient region, such as its area, shape, and location, will participate in the judgment logic of the candidate box selection process.
[0035] Defect response parameters are generated by filtering candidate boxes based on salient regions. The minimum bounding rectangle of each salient region is calculated as the initial candidate box, and the center coordinates, width, and height of the bounding rectangle are directly extracted from the geometric properties of the salient region. Connected components with an area less than 50 pixels in a salient region are considered noise responses and are discarded. Normal processing textures on the surface of metal workpieces may occasionally produce slight spurious responses, which are filtered out using an area threshold. Multiple adjacent and overlapping salient regions are merged when the overlapping area of the bounding rectangles of the two regions exceeds 30% of the area of the smaller rectangle. Long scratches on the surface of metal workpieces that are divided into two segments due to local color differences are merged to generate an expanded candidate box that encloses the complete scratch. The defect response parameters include five fields: center coordinates, width, height, confidence score, and the location index of the corresponding feature map. The confidence score is obtained by normalizing the average response intensity within the corresponding salient region. Areas with severe surface corrosion receive higher confidence scores due to abnormally significant textures, while areas with early slight oxidation receive relatively lower confidence scores due to weak color differences. The defect response parameters are organized in a list, and the list length is equal to the number of candidate boxes obtained from the current image.
[0036] Step S130: Spatial pyramid pooling is performed on the feature map to determine the fused features. The fused features are used to perform bounding box regression to locate the defect segment. Based on the defect segment features, the defect response parameters are fused to form a feature matrix.
[0037] Specifically, spatial pyramid pooling is used to determine the fusion features of the feature map. During spatial pyramid pooling, the feature map is divided into four granularities: 1×1, 2×2, 4×4, and 8×8. The feature vector within each grid cell is compressed into a single representative vector through max pooling. Contextual relationships across different spatial scales affect the defect recognition effect. Rolling defects on the surface of metal workpieces often exhibit a striped distribution along the rolling direction. Spatial pyramid pooling performs region division and pooling aggregation at multiple scales to capture directional correlation information spanning a large spatial range. After 1×1 grid pooling, the feature map generates one global feature vector, which captures macroscopic information such as the overall statistical characteristics of the image, including average brightness and global texture density. After 8×8 grid pooling, 64 local feature vectors are generated, which preserve the detailed differences between regions. The fusion feature is formed by concatenating the pooling results of the four granularities along the channel dimension. The dimension of the fusion feature is equal to the original number of channels multiplied by the total number of grid cells, which is 1+4+16+64=85 cells, containing multi-level spatial context information from global to local. Two defect regions that are far apart are assigned to the same grid cell at the coarse-grained level of the spatial pyramid, and their feature correlation is established. Multiple pore defects scattered on the surface of the casting exhibit an aggregated response pattern in the fusion feature.
[0038] In some embodiments, the step of using the fusion features to locate the defect segment by bounding box regression includes: generating a location prediction map by mapping the fusion features to the detection space; generating an initial bounding box by generating anchor boxes from the location prediction map; extracting defect morphological features from the initial bounding box and performing aspect ratio constraint regression to form a morphological correction box; and performing non-maximum suppression on the morphological correction box to form the defect segment.
[0039] A location prediction map is generated by mapping the fused features to the detection space. The fused features undergo a 1×1 convolution transformation to generate the location prediction map. This transformation compresses the high-dimensional vector of each feature location in the fused features into a low-dimensional location code, which includes the probability of a defective target existing at that location and the predicted bounding box parameters. There is a scale difference between the feature space and the pixel space of the original image. A location in the feature space corresponds to a receptive field region in the original image, rather than a single pixel. The size of the receptive field region depends on the downsampling factor of the network; with a 4x downsampling, a location in the feature map corresponds to a 4×4 pixel region in the original image. Bounding box regression maps the response of the fused features in the feature space to coordinates in the pixel space. Locations with higher response intensity exhibit a higher probability of target presence in the location prediction map. For example, the peeling defect region on the surface of a metal workpiece forms a significant high-probability response peak in the location prediction map. The spatial resolution of the location prediction map is consistent with the feature map resolution of the fused features. Each spatial location outputs a set of predicted values, including parameters such as foreground probability, center offset, and size scaling factor. The mapped location prediction map retains the contextual information of the multi-scale spatial pyramid while transforming abstract features into location parameters with clear geometric meaning. Locations with a foreground probability below 0.5 in the location prediction map are identified as background areas and filtered out, while only high-probability locations are retained for subsequent anchor frame generation.
[0040] An initial bounding box is obtained by generating anchor frames from the location prediction map. Each high-probability location in the location prediction map is associated with a set of preset anchor frame templates. The anchor frame templates define the reference shape and size of the bounding box of the defect that may appear at that location. The initial bounding box is generated by combining the anchor frame templates with the offset parameters in the location prediction map. The anchor frame templates provide the reference position and size, and the offset parameters fine-tune the reference values to fit the boundary of the actual defect. The anchor frame template design covers a variety of aspect ratios and scale combinations. The aspect ratios are set to 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, and 4:1, a total of 7 specifications to accommodate different defect shapes. Each aspect ratio is configured with 3 scales: small, medium, and large. The side length of the small-scale anchor frame is 32 pixels, the medium-scale is 64 pixels, and the large-scale is 128 pixels. A total of 21 anchor frame candidates are generated for a single location. Point-like pore defects on the surface of castings are close to circular and are adapted to 1:1 anchor frames. Linear scratches on the surface of metal workpieces are thin and long and are adapted to 1:3 or 1:4 anchor frames. Multiple initial bounding box candidates of different sizes are generated for each high-probability location. The number of candidate boxes is equal to the number of anchor box templates associated with that location. The coordinate parameters of the initial bounding boxes are calculated by overlaying the offset prediction values from the location prediction map onto the anchor box template coordinates. The unit of the center point translation in the offset prediction values is pixels, and the size scaling is a proportional coefficient relative to the anchor box size. Multiple adjacent high-probability locations may correspond to the same defect target, and the initial bounding boxes generated at each location have a large overlap. These redundant boxes will be merged in the subsequent suppression process.
[0041] The initial bounding box extracts defect morphological features, and aspect ratio-constrained regression is used to form a morphological correction bounding box. The image region defined by the initial bounding box is used to extract feature vectors through region-of-interest pooling. Pooling maps bounding box regions of different sizes to a fixed-size 7×7 feature map. After flattening, the feature map forms a feature vector that is input into the morphological regression network. The initial bounding box generated based on a preset anchor frame template may deviate from the actual contour of the defect. Folded defects on the surface of metal workpieces actually present as curved stripes, while a rectangular bounding box can only provide a rough envelope. Aspect ratio-constrained regression fine-tunes the shape parameters of the initial bounding box based on the defect morphological features within the bounding box. The morphological regression network contains two fully connected layers: the first layer compresses the feature vector dimension from 7×7×channels to 256 dimensions, and the second layer outputs a 4-dimensional refined parameter vector. The refinement parameters include an aspect ratio adjustment factor and a size correction factor. The aspect ratio adjustment factor ranges from 0.5 to 2.0, indicating that the aspect ratio can be adjusted to 0.5 to 2 times the original. The size correction factor ranges from 0.8 to 1.2, indicating that the area can be adjusted to 80% to 120% of the original. The aspect ratio constraint ensures that the aspect ratio of the refined bounding box matches the actual shape of the defect within the box, avoiding mismatches such as wide, flat defects with tall, narrow boxes or slender defects with square boxes. The shape-corrected box is formed after the aspect ratio constraint regression is completed. Compared to the initial bounding box, the shape-corrected box has shape parameters that better fit the actual defect contour. After shape correction, the aspect ratio of the bounding box for linear defects on the metal workpiece surface is more consistent with the defect direction. Candidate boxes with large shape deviations show significantly improved positioning accuracy after correction, while candidate boxes with small shape deviations show limited changes.
[0042] Non-maximum suppression (NMS) is applied to the morphological correction boxes to form defect segments. The morphological correction boxes are sorted in descending order of confidence score, and are processed sequentially starting with the candidate box with the highest score. For each candidate box processed, the Intersection over Union (IoU) with all subsequent candidate boxes is calculated. Multiple highly overlapping candidate boxes may exist for the same defect target. These redundant boxes originate from repeated detections of the same target at adjacent feature locations and multiple matches of the same target by anchor boxes of different sizes. Without suppression, a single defect will be counted repeatedly. NMS uses IoU to measure the degree of overlap: IoU = intersection area of two boxes / union area of two boxes. Subsequent boxes with an IoU exceeding 0.5 are considered redundant and discarded. An IoU threshold of 0.5 indicates that two boxes with an overlap area exceeding half of the union area are considered to point to the same target. The same wear defect on the surface of a metal workpiece may generate multiple overlapping morphological correction boxes. NMS retains the box with the highest score as the final location result for the defect, and discards the remaining overlapping boxes. Defect segments are formed after non-maximum suppression. Each defect segment corresponds to an independent defect target, and the defect segments are independent of each other and do not have high overlap. The number of shape-corrected boxes is significantly reduced after suppression. Typically, the number of defect segments after suppression is about 1 / 5 to 1 / 3 of the number of shape-corrected boxes before suppression. After suppression, 40 to 70 defect segments are usually retained after processing the first 200 candidate boxes. Defect segments are output in the form of bounding box coordinates and confidence scores. The coordinate parameters include four values: the top-left and bottom-right horizontal and vertical coordinates, plus the confidence score, for a total of five parameters that fully describe a defect segment.
[0043] A feature matrix is formed by fusing defect response parameters with defect segment features. The image region defined by each defect segment is mapped to a fixed-dimensional feature vector through region-of-interest pooling. The pooled feature vector encodes the visual attributes of the defect, such as texture pattern, shape contour, and color distribution. Location information alone is insufficient for defect type identification and severity assessment. The feature representation within the defect segment is fused with the response parameters generated in the preceding steps to construct a complete defect description. The confidence score, location index, and other information recorded in the defect response parameters are concatenated with the defect segment features. The concatenation method involves linking the feature vector with the five field values of the defect response parameters to form an extended feature vector. The extended feature dimension equals the original feature dimension plus 5. The confidence score reflects the probability that the region contains a real defect; incorporating it into the feature representation helps distinguish between real defects and false positives. The feature matrix is formed after feature fusion. The number of rows in the matrix equals the number of defect segments, i.e., the number of detected candidate defects. The number of columns in the matrix equals the fused feature dimension, which is typically 256 or 512 dimensions plus 5 dimensions of response parameters, totaling 261 or 517 dimensions. Different types of defects on the surface of a metal workpiece correspond to different row vectors in the feature matrix. The feature vectors of different defects exhibit a separable clustered distribution in high-dimensional space. Each row of feature vectors in the feature matrix fully describes the visual attributes and response characteristics of a candidate defect, forming the basis of input data for defect classification and rating. The location index field in the defect response parameters is retained in the feature matrix, recording the pixel coordinate range of the defect in the original image, and the correspondence between the feature vector and the location in the original image is traceable.
[0044] Step S140: Perform pattern clustering on the feature matrix to determine the category distribution, extract weak response regions from the feature matrix to generate early warning markers, generate defect level parameters based on the category distribution and early warning markers, and extract severity gradient features from the defect level parameters to construct a defect feature map.
[0045] In some embodiments, the step of determining the category distribution by pattern clustering of the feature matrix includes: generating a sample distance metric based on the feature matrix; extracting outlier samples from the sample distance metric to generate a boundary correction factor; initializing centroids using the sample distance metric and combining the boundary correction factor to form dynamic cluster centers; and classifying samples according to the dynamic cluster centers to form a category distribution.
[0046] A sample distance metric is generated based on the feature matrix. The pairwise distances between all row vectors in the feature matrix are calculated using Euclidean distance, which measures the absolute difference between feature vectors and is suitable for features with consistent scale. Cosine distance measures the directional difference between feature vectors and is suitable for features with large amplitude variations. Each row vector in the feature matrix represents a different defect sample. Cluster analysis quantifies the similarity between any two samples, measured by distance in the feature space. A smaller distance indicates greater similarity between the two samples and a higher probability that they belong to the same type of defect. When the feature matrix contains N defect samples, the sample distance metric forms an N×N symmetric matrix, where the element in the i-th row and j-th column represents the distance between the i-th and j-th samples. The eigenvectors of porosity and micro-porosity defects on the surface of castings show relatively small differences, and the sample distance metric between these two types of defect samples is typically below 0.3. Conversely, the eigenvectors of porosity defects and scratch defects show significant differences, and the sample distance metric between these two types of samples is typically above 0.7. In the sample distance metric matrix, diagonal elements with values of zero represent the distance between a sample and itself, while off-diagonal elements reflect the degree of difference between samples. When the feature dimension in the feature matrix is high, exceeding 256, the computational complexity of the sample distance metric increases accordingly. Efficiency can be improved through PCA dimensionality reduction preprocessing or the approximate nearest neighbor method.
[0047] Outlier sample boundary correction factors are generated from the sample distance metric. The average distance of each sample is calculated by the average distance of that sample to all other samples. Samples with an average distance exceeding twice the global mean are identified as outliers. Some samples in the sample distance metric matrix have large distances to all other samples. These samples exist in isolation in the feature space, far from any clustering region, and are called outliers or anomalous samples. The causes of outliers include rare defect types, detection noise interference, false detection areas, etc. Including outliers in regular clustering will shift the cluster centers and affect the clustering quality. Boundary correction factors are generated for outliers identified in the sample distance metric. The factor's role is to reduce the influence weight of outliers on cluster boundaries during the clustering process. Occasional foreign matter contamination on the surface of metal workpieces has a feature pattern in the feature matrix that differs significantly from common defects such as scratches and pores. These contamination samples are isolated in the sample distance metric and are therefore identified as outliers. Boundary correction factors are represented as a pair of sample indices and correction coefficients. The correction coefficient for outliers is set to a small value between 0.1 and 0.5. The correction coefficient is negatively correlated with the average distance; the larger the average distance, the smaller the correction coefficient. The correction coefficient for non-outliers is 1.0, indicating no correction. The proportion of outliers in the sample distance metric typically does not exceed 5% of the total sample size. If the proportion exceeds 10%, it suggests a potential problem in the feature extraction process and should be investigated.
[0048] Dynamic cluster centers are formed by initializing centroids using a sample distance metric combined with a boundary correction factor. The density distribution information in the sample distance metric is obtained by calculating the local density estimate for each sample. Local density is defined as the number of samples within a radius *r* centered on that sample, where *r* is set to 0.5 times the mean of the sample distance metric. The selection of initial centroids in the clustering algorithm affects the final clustering result. Randomly selecting initial centroids may lead to local optima or slow iterative convergence. Intelligent initialization based on sample distribution characteristics can improve convergence. The K positions with the highest density in the sample distance metric are selected as the initial centroids for the K clusters. The value of K is determined by the silhouette coefficient, S = (ba) / max(a,b), where *a* is the average distance from the sample to other samples in the same cluster, and *b* is the average distance from the sample to the nearest sample in a different cluster. The largest silhouette coefficient within the range of 2 to 10 is selected for K. The boundary correction factor participates in the weighted calculation during the centroid initialization stage. Outliers have a lower correction coefficient in the boundary correction factor, weakening their contribution to density estimation. The introduction of the boundary correction factor avoids the misselection of low-density regions near outliers as cluster centers. Common defects on the surface of metal workpieces include scratches, peeling, and cracks. Each type of defect sample forms an independent high-density cluster in the sample distance metric. During dynamic cluster center initialization, a center point is selected in each cluster. After initialization, the dynamic cluster center enters an iterative optimization phase. In each iteration, the center coordinates of each cluster are recalculated based on the sample affiliation. The iteration terminates when the center point displacement is less than 0.01 or the number of iterations reaches 100.
[0049] The sample classification is performed based on dynamic cluster centers to form a category distribution. The distance between each row vector in the feature matrix and all dynamic cluster centers is calculated one by one, and the index of the center with the smallest distance is used as the cluster label for that sample. After the dynamic cluster centers are determined, each sample in the feature matrix is assigned to a specific cluster based on its distance from each cluster center. For example, a sample with porosity defects on the surface of a casting is closest to the dynamic cluster center of the porosity cluster, and its cluster label is set to the porosity class; a sample with crack defects on the same workpiece is closest to the dynamic cluster center of the crack cluster, and its cluster label is set to the crack class. The category distribution is formed after all samples have been classified. The category distribution stores the cluster labels of each sample in vector form, with the vector length equal to the number of rows in the feature matrix. The dynamic cluster centers may shift after iterative optimization, causing changes in the classification of some boundary samples. The category distribution is recalculated after each iteration until it stabilizes. The category distribution also tracks the sample quantity distribution of each cluster. In metal workpiece inspection, if the number of samples of a certain type of defect increases abnormally, it indicates that there may be a deviation in the processing parameters. The dynamic cluster center coordinates are output synchronously after the category distribution stabilizes. The center coordinates are stored in the form of a K×D matrix, where K is the number of clusters and D is the feature dimension. The center coordinates represent the typical feature patterns of each type of defect and can be used for rapid classification of new samples.
[0050] In some embodiments, the step of extracting weak response regions from the feature matrix to generate early warning markers includes: generating a response intensity distribution based on the feature matrix; identifying response decay trends from the response intensity distribution to locate early defect formation regions; performing time-series stability verification on the early defect formation regions to filter effective weak response regions; and generating early warning markers based on the effective weak response regions.
[0051] The response intensity distribution is generated based on the feature matrix. The comprehensive response intensity is calculated row by row in the feature matrix by taking the absolute values of each component of the vector in each row and then summing them with weights. The weights are initially set to be equal, i.e., each dimension has a weight of 1 / D, where D is the feature dimension. The component values of each row vector in the feature matrix reflect the activation intensity of the corresponding defect sample in each feature dimension. The overall response intensity of the sample is obtained by aggregating the component values. A high response intensity indicates that the defect features are significant and easily identifiable, while a low response intensity indicates that the defect features are weak and may be in the early nascent stage. The response intensity distribution is formed by calculating and summing the comprehensive response intensity row by row in the feature matrix. In the feature matrix, the component values of the feature vectors of mature crack samples are generally large, and the comprehensive response intensity is in the high-value range of the distribution; the component values of the feature vectors of early crack initiation areas are generally low, and the comprehensive response intensity is in the low-value range of the distribution. The response intensity distribution is represented as a histogram or probability density curve, with the horizontal axis representing the response intensity value and the vertical axis representing the number of samples or probability density in the corresponding intensity range. In fatigue testing of metal workpiece surfaces, the response intensity distribution exhibits a bimodal characteristic. The peak is located in the response intensity range of 0.7 to 0.9, corresponding to established mature defects, while the trough is located in the response intensity range of 0.2 to 0.4, corresponding to early-stage defect regions still in their nascent stage. Statistical characteristics of the response intensity distribution, such as mean, standard deviation, and kurtosis, are calculated simultaneously. These indicators are used to determine the threshold for identifying weak response regions.
[0052] Early defect formation areas are located by identifying response decay trends from the response intensity distribution. Samples with response intensities lower than 0.5 times the global mean are extracted as weak response samples. The feature vectors of weak response samples are compared with the cosine similarity of mature defect samples of the same type. Samples in the low-value range of the response intensity distribution are further analyzed to determine whether they belong to early defects rather than simply false positives. The analysis is based on whether the feature patterns of weak response samples conform to the early evolution characteristics of a certain type of defect. Response decay trend identification compares the features of weak response samples in the response intensity distribution with those of mature defect samples of the same type. If the feature vector direction of the weak response sample is consistent with that of the mature defect sample but the amplitude decays proportionally, and the cosine similarity is higher than 0.8 and the amplitude ratio is lower than 0.5, then the weak response sample is determined to be an early form of that type of defect. Early defect formation areas are located from weak response samples that conform to the response decay trend. Each early defect formation area corresponds to a suspected emerging defect target. Early fatigue damage in metal workpieces manifests as the initiation of surface microcracks. This early damage is located in the low-value range of the response intensity distribution, and its eigenvector shows an attenuating relationship with mature crack defects, thus identifying it as an early defect formation region. Samples with extremely low response intensity and eigenmodes that do not match any known defect type are considered false positives and excluded as early defect formation regions. Early defect formation regions are recorded using a pairing of sample index and suspected defect type. The index is used to associate the original image location, and the type guides targeted re-inspection strategies.
[0053] For example, the step of performing temporal stability verification on the early defect formation area to screen effective weak response regions includes: establishing a multi-frame response correlation based on the early defect formation area; performing response fluctuation detection along the multi-frame response correlation to form a stability index; performing transient noise interference analysis through the stability index to determine a persistent low response region; and performing regional validity calibration based on the persistent low response region to form an effective weak response region.
[0054] A multi-frame response correlation is established based on the early defect formation area. The corresponding positions of the early defect formation area in continuously acquired multi-frame images are registered and tracked using feature point matching or template correlation methods. The response intensity of the corresponding positions in each frame is extracted and summarized to form a response time sequence. While the early defect formation area is located in a single-frame image, a single-frame image is insufficient to distinguish between persistent real defect signals and transient interference signals. Time-series analysis of multi-frame image sequences can differentiate between these two types of signals. The multi-frame response correlation is organized in a paired manner using position indices and response sequences. The position index identifies the spatial location of the early defect formation area in the image, and the response sequence records the response intensity values of that location in each frame. Four to eight frames of images are continuously acquired at the inspection station. The position coordinates of the early defect formation area in the first frame are mapped to the corresponding positions in the remaining frames after registration transformation. In metal workpiece surface inspection, the workpiece moves through the camera's field of view at a constant speed, ranging from 0.5 to 2 meters per second. With a frame interval of 50 milliseconds, the workpiece displacement ranges from 25 to 100 millimeters. The early defect formation area translates along the direction of motion in consecutive frame images. Multi-frame response correlation is established considering motion compensation to ensure that the same physical location is being tracked. Some areas within the early defect formation area may be invisible in certain frame images due to workpiece edges or occlusion; these incomplete response sequences are marked as missing values in the multi-frame response correlation.
[0055] Stability indices are generated by performing response fluctuation detection along multi-frame response correlation. The coefficient of variation (COP) of each response sequence is calculated, which is equal to the standard deviation of the response sequence divided by the mean. The smaller the COP, the more stable the response; a COP below 0.3 is considered stable. The response sequence at each location in the multi-frame response correlation reflects the change in response intensity over time at that location. The response intensity of the true early defect region is relatively stable with small fluctuations, while the response intensity of the false weak response region caused by transient interference changes drastically. The response intensity of early folding defects on the surface of metal workpieces remains stable across frames because the material has undergone substantial deformation, and the COP is usually below 0.2, resulting in a small COP calculated for the stability index. In contrast, the position and shape of transient occlusion caused by oil splashes change significantly across different frames, and the COP is usually above 0.5, resulting in a large COP calculated for the stability index. For locations with missing values in the response sequence of the multi-frame response correlation, the stability index is calculated only based on valid frames. If the number of valid frames is less than 3, the location is marked as unassessable for stability evaluation. The stability index also includes a response trend component. This trend component is obtained by linearly fitting the response sequence to obtain the slope. A positive slope indicates that the response intensity is increasing over time, potentially reflecting defect development, while a negative slope indicates that the response intensity is decreasing over time, potentially reflecting the fading of transient interference. The stability index is represented by two values: the coefficient of variation and the trend slope. These two values are used together for subsequent noise interference detection.
[0056] Persistent low-response regions were identified through transient noise interference analysis using stability indices. Regions with a coefficient of variation below 0.3 and an absolute value of the trend slope below 0.1 were classified as persistent defects, while regions with a coefficient of variation above 0.3 or a significantly negative trend slope (below -0.2) were classified as transient interference. The coefficient of variation and trend slope components of the stability indices jointly characterize the temporal features of the response sequence, thereby distinguishing between genuine persistent defects and spurious transient interference in each early defect formation area. Persistent low-response regions were identified from early defect formation areas that met the criteria; these regions exhibited stable low-response characteristics rather than random fluctuations in time. Early fatigue damage on the surface of metal workpieces exhibited a low coefficient of variation and a near-zero trend slope in the stability indices, thus these regions were identified as persistent low-response regions. Conversely, abnormal reflections caused by uneven surface oil film exhibited a high coefficient of variation in the stability indices, thus these regions were classified as transient interference and excluded. Although regions with a significantly positive trend slope in the stability index are not transient disturbances, their increasing response intensity may reflect rapid defect development. Regions with a trend slope higher than 0.2 are separately marked as developing defects rather than stable, persistently low-response regions. The number of persistently low-response regions is less than the number of early defect formation regions; the difference represents regions identified as transient disturbances or developing defects.
[0057] Effective weak response regions are formed by calibrating the validity of persistently low-response regions. The cosine similarity between the feature vector of a persistently low-response region and the known defect template vector is calculated for each region. Regions with a similarity exceeding 0.6 are calibrated as effective weak response regions. While transient interference has been eliminated from the evaluation of persistently low-response regions, their practical value for defect early warning is still further verified. Some persistently low responses may be caused by normal surface texture features such as machining marks or grinding traces, rather than defect initiation. Region validity calibration assesses the defect relevance of persistently low-response regions based on the degree of matching between the region's feature pattern and the known defect template. This matching degree is quantified by calculating the cosine similarity between the feature vector and the defect template vector. Effective weak response regions are formed by selecting regions from the persistently low-response regions that meet the defect relevance criteria. Early defect initiation regions on the metal workpiece surface were identified in the persistently low-response regions, where the similarity between their feature vectors and the defect templates exceeded 0.6. These regions were selected as valid weak-response regions. While normal grinding textures on the workpiece surface also exhibited persistently low responses, their feature vectors had similarities below 0.6 with any defect template; these regions were not included in the valid weak-response regions. Within the persistently low-response regions, some areas did not perfectly match the existing defect templates nor conform to normal texture features; regions with similarities between 0.4 and 0.6 were marked as unknown-type weak-response regions and included in the valid weak-response regions for manual review.
[0058] Early warning markers are generated based on effective weak response regions. For each effective weak response region, three fields—warning type, warning level, and warning location—are filled to generate the early warning marker. The warning type corresponds to the suspected defect category, such as early cracking, early corrosion, or early spalling. Effective weak response regions, verified through time-series stability testing, are confirmed as genuine early defect initiation areas rather than transient interference. These regions are assigned clear warning labels to attract the attention of quality control personnel. The warning level reflects the urgency of developing into a mature defect and is divided into three levels: low, medium, and high. The ratio of the response intensity of the effective weak response region to the response intensity of similar mature defects is used to determine the warning level. A ratio below 0.3 indicates a low warning level, a ratio between 0.3 and 0.5 indicates a medium warning level, and a ratio above 0.5 indicates a high warning level, signifying that it is about to evolve into a mature defect. The warning location records the position coordinates of the region in the workpiece coordinate system. The coordinate format is (x_min, y_min, x_max, y_max), representing the outer rectangular boundary of the warning region. Early cracks on the surface of metal workpieces are identified in the effective weak response area. An early warning marker indicates that this area is a high-level crack warning, suggesting that the workpiece should be closely monitored or addressed in advance to prevent further defect development. The early warning marker is independent of the category distribution determined in the preceding steps. The category distribution classifies mature defects, while the early warning marker warns of potential defects that are still developing. Together, they constitute a complete defect status assessment system.
[0059] Defect severity parameters are generated based on category distribution and early warning markers. The category labels in the category distribution and the warning levels in the early warning markers jointly determine the defect severity parameters. Mature defects without warning markers are classified into levels based on defect type and size, while early defects with warning markers are classified into levels based on warning level and development trend. The category distribution reveals the type category to which each defect sample belongs, and the early warning markers identify potential early defect risk areas. These two types of information are integrated to form a comprehensive assessment of defect severity. Defects of the same type have different severity levels depending on their development stage; early-stage defects have lower severity levels than fully developed defects. Porosity defects on the surface of castings are classified as pores in the category distribution. If the pore diameter exceeds the allowable upper limit of 2 mm, the defect severity parameter is marked as severe; if the diameter is within the allowable range, it is marked as minor. If the same pore is marked as an expanding risk area in the early warning markers, the defect severity parameter is increased by one level based on the size level. Different grading standards are applied to defect samples belonging to different clusters in the category distribution. Scratch defects are mainly graded based on length (more than 10 mm is severe) and depth (more than 0.1 mm is severe), while rust defects are mainly graded based on area (more than 25 square millimeters is severe) and degree of corrosion. Defect grade parameters are output simultaneously in two forms: grade label and grade score. The grade label is a discrete four-level classification of severe / moderate / minor / warning, and the grade score is a continuous value in the range of 0 to 100 to facilitate fine-grained sorting.
[0060] Severity gradient features are extracted from defect level parameters to construct a defect feature map. The defect level parameters are mapped to a two-dimensional coordinate system on the workpiece surface. Three types of information are extracted to form the severity gradient features: the increasing trend of severity along a specific direction, the spatial clustering of severity, and the location distribution of severity extreme points. While the defect level parameters provide the level assessment results for each defect sample, isolated level values do not reveal the spatial distribution pattern and severity trend of defects on the workpiece surface. The level information is fused with spatial location information to construct a visualized defect distribution map. The severity gradient features describe the spatial variation of the defect level parameters. If high-level defects in the defect level parameters are concentrated in a certain area of the workpiece, then there is a systemic quality problem in that area. If crack defects on the surface of a metal workpiece show an increasing severity trend along a certain direction (i.e., a gradient value greater than 0.1), it indicates a deviation in the processing technology. The defect feature map is constructed by mapping defect level parameters to a two-dimensional coordinate system on the workpiece surface. The map displays the spatial distribution of severity in the form of a heatmap, using 256 levels of pseudo-color coding. Severity scores of 0-25 correspond to blue tones, 26-50 to green tones, 51-75 to yellow tones, and 76-100 to red tones. High-level defect areas are displayed as dark, high-temperature areas, low-level defect areas as light, low-temperature areas, and defect-free areas are displayed as the background color. Warning areas in the defect level parameters are marked with dashed boundaries in the defect feature map, distinguishing them from the solid boundaries of existing defects.
[0061] Step S150: Feature sensitivity weights are extracted by performing feature optimization on the defect response parameters and feature matrix through multi-scale loss constraints. Weighted discrimination rules are constructed based on the feature sensitivity weights. The defect feature map is then integrated through the weighted discrimination rules to output the visual detection results.
[0062] In some embodiments, the step of extracting feature sensitivity weights by performing feature optimization on the defect response parameters and the feature matrix through multi-scale loss constraints includes: performing multi-scale decomposition on the defect response parameters and the feature matrix to form hierarchical loss values; counting the number of samples of each category from the feature matrix to generate class balance weights; adaptively weighting the hierarchical loss values with the class balance weights to form a balance loss; and performing gradient updates based on the balance loss to form feature sensitivity weights.
[0063] Multi-scale decomposition of the defect response parameters and feature matrix yields hierarchical loss values. The defect response parameters and feature matrix are downsampled sequentially to three scale levels: 1 / 2, 1 / 4, and 1 / 8. The deviation between the predicted value and the true label is calculated at each scale level to obtain the hierarchical loss value. The defect response parameters focus on the overall response intensity at the region level, while the feature matrix focuses on local feature patterns at the pixel level. Joint optimization within the multi-scale framework unifies the scale differences between the two types of data. In defect detection of metal workpieces, coarse-scale loss focuses on whether there are obvious defect areas on the workpiece surface, while fine-scale loss focuses on whether there are subtle feature differences within the defect areas. Multi-scale joint constraints enable the detection model to possess both macroscopic judgment and microscopic perception capabilities. The difference between the confidence score and the true label in the defect response parameters is used to calculate cross-entropy loss at each scale level, and the difference between the feature vector and the target feature in the feature matrix is used to calculate mean squared error loss at each scale level. The sum of these two types of losses forms the comprehensive loss for each scale. Hierarchical loss values are stored as vectors of length 3, corresponding to 3 scale levels. The k-th element in the vector represents the loss value at the k-th scale level. The magnitudes of the loss values at each scale may differ. Typical values for coarse-scale loss are 0.1 to 0.3, for meso-scale loss are 0.3 to 0.6, and for fine-scale loss are 0.5 to 1.0. Directly adding them together may cause the fine-scale loss to dominate the optimization direction while ignoring the coarse-scale information.
[0064] The class balance weights are generated by statistically analyzing the number of samples in each category from the feature matrix. The sample counts for each category in the feature matrix are grouped and statistically analyzed. The inverse frequency weights are obtained using the formula w_c = (N_total / N_c) / Σ(N_total / N_i), where w_c is the class balance weight for the c-th defect, N_total is the total number of samples, N_c is the number of samples in the c-th category, N_i is the number of samples in the i-th category, i iterates through all defect categories, and Σ(N_total / N_i) is the sum of the inverse frequencies of each category used for normalization. Categories with fewer samples typically have a larger weight. There is often a significant imbalance in the number of samples for different types of defects; common defects such as scratches and stains have far more samples than rare defects such as cracks and pores. This class imbalance causes the optimization process to favor the majority class. In surface inspection of coated parts, common defect samples typically account for 60% to 80% of the total, while rare defect samples usually account for only 5% to 15%. Category balancing weights assign greater weights to rare defects to improve the model's sensitivity in identifying these types of defects. The category label of each sample can be obtained from the category distribution in the feature matrix through the preceding steps. Grouping and statistically analyzing by category label yields the number of samples in each category. The category balancing weights are stored as a vector indexed by the category label, with a vector length equal to the total number of defect categories. The c-th element in the vector represents the balancing weight of the c-th defect category. The weight of extremely rare categories in the category balancing weights is set to an upper limit of 10 to avoid excessive weights due to insufficient sample size, which could lead to instability in the optimization process.
[0065] A balanced loss is formed by adaptively weighting the hierarchical loss values using category balancing weights. The loss at each scale in the hierarchical loss values is weighted according to a scale weight coefficient, set at 0.2 for coarse scale, 0.3 for medium scale, and 0.5 for fine scale. The fine scale has a larger weight to emphasize the learning of detailed features. The determination of the fusion weights considers both the relative importance of each scale and the balancing needs of each category. The category balancing weights introduce category dimension adjustment during the fusion process to correct optimization biases caused by category imbalance. The balanced loss is obtained by weighting and summing the losses at each scale in the hierarchical loss values according to the category balancing weights, and then taking the average. The weighting method is to multiply the loss of each sample by the balancing weight of its category, and then average the weighted losses of all samples to obtain the batch-level balanced loss. In the detection of surface defects in metal workpieces, crack defects are few in number but cause serious damage. The category balancing weights assign a large weight to the crack class, typically 3 to 5 times that of the scratch class. Misclassification of crack samples contributes significantly to the balanced loss, guiding the optimization process to focus on the accuracy of crack defect identification. As the final form of the optimization objective function, the decrease in the value of the balanced loss indicates that the model's defect identification ability is improving. The introduction of class balance weights enables the balanced loss to give fair attention to all types of defects, and the model has a good ability to identify both common and rare defects.
[0066] Feature sensitivity weights are formed by gradient updates based on balanced loss. The partial derivatives of the balanced loss with respect to each column of the feature matrix (i.e., each feature dimension) are obtained through backpropagation. The absolute values of these partial derivatives are normalized and then converted into feature sensitivity weights. The difference in contribution of each feature dimension to the loss reduction during parameter updates reflects the sensitivity of that dimension. A larger absolute gradient value indicates a more significant impact of that feature dimension on the loss, a higher contribution to defect identification, and a correspondingly larger sensitivity weight. In metal workpiece inspection, edge and texture features of defects contribute significantly to defect identification. The gradients of the balanced loss with respect to these two feature dimensions are large, typically 2 to 3 times that of other dimensions, resulting in higher weights for these dimensions in the feature sensitivity weights. Conversely, the texture features of the workpiece background contribute very little to defect identification, with gradients for these dimensions approaching zero, resulting in lower weights for these dimensions. The feature sensitivity weights tend to stabilize after multiple iterations of gradient updates, typically 50 to 100 iterations. The stable weight distribution reflects the final ranking of the contribution of each feature dimension to the current detection task. The feature sensitivity weights are extracted and output after the balance loss converges. The dimension of the weight vector is the same as the number of columns of the feature matrix. The weight values of each dimension are normalized to the range of 0 to 1 to facilitate fusion with the discrimination rules.
[0067] A weighted discrimination rule is constructed based on feature sensitivity weights. The feature sensitivity weights, as feature weighting coefficients, are embedded in the classification decision function. The decision function calculates the matching score S_c according to the formula S_c=Σ(w_i×f_i×t_ci), where w_i is the i-th dimension feature sensitivity weight, f_i is the i-th dimension of the input feature vector, and t_ci is the i-th dimension of the c-th class template. The class with the highest matching score is used as the predicted class. In the detection of surface defects in castings, edge response features contribute significantly to the identification of cracks and porosity, hence the corresponding dimension in the feature sensitivity weights has a larger value. The edge response dimension weight is typically 0.7 to 0.9. Color response features contribute significantly to the identification of oxidation discoloration, hence the corresponding dimension value is also large. The weighted discrimination rule assigns different levels of attention to different features based on these differentiated weights during classification. The weighted discrimination rule comprises two components: a category determination module and a confidence assessment module. The category determination module outputs the predicted defect type label, while the confidence assessment module outputs the confidence score of this prediction. The confidence score is represented by the normalized value of the difference between the highest and second-highest matching scores. Dimensions with a feature sensitivity weight below 0.1 are considered invalid features and are masked. This feature filtering mechanism reduces the interference of noisy features on the discrimination results. The discrimination threshold in the weighted discrimination rule is set to 0.5. Match scores above 0.5 are judged as defects, and scores below 0.5 are judged as non-defects. The threshold parameter is adjusted according to the tolerance for false negatives and false positives in the detection scenario.
[0068] The visual inspection results are output by integrating the defect feature map using weighted discrimination rules. The feature vector of each defect region marked in the defect feature map is extracted and input into the weighted discrimination rules. The type prediction and confidence assessment results are written back to the corresponding positions in the defect feature map to form complete annotations. The weighted discrimination rules provide the mapping logic from features to categories, while the defect feature map provides information on the spatial distribution and severity gradient of defects on the workpiece surface. The integration of these two elements forms a complete inspection output containing full-dimensional information such as the location, type, level, and distribution of defects. In the inspection of metal workpiece surfaces, the defect feature map marks several high-response areas and warning areas. The weighted discrimination rules classify high-response areas as scratches, porosity, corrosion, etc., and warning areas as early-stage cracks, etc. The visual inspection results are organized in a structured report format, including four modules: inspection image, defect annotation map, defect list, and statistical summary. The heatmap visualization information from the defect feature map is retained in the defect annotation module of the visual inspection results. The resolution of the defect annotation map is consistent with the original image, and the annotation colors are distinguished according to the defect level, making it easy for quality inspectors to intuitively grasp the defect distribution. The defect list module in the visual inspection results records detailed information for each defect, with each record containing six fields: defect number, location coordinates, defect type, severity level, confidence score, and area size. If the visual inspection results of a metal workpiece show multiple defects on the surface and the severity increases along a certain direction, it indicates an abnormality in the processing technology, and relevant parameters should be investigated. The statistical summary module in the visual inspection results summarizes and analyzes the defect distribution of the current batch of workpieces. The summary content includes statistical indicators such as the proportion of various types of defects, average severity score, and hot spots in defect density distribution, providing data support for production process improvement.
[0069] To implement the machine vision-based workpiece defect detection method corresponding to the above method embodiments, and to achieve the corresponding functions and technical effects. See also Figure 3 , Figure 3 This paper illustrates a structural block diagram of a machine vision-based workpiece defect detection system 300 according to an embodiment of this application, including: Image acquisition module 301 is used to acquire image sequence data of the workpiece surface, and perform size normalization and data enhancement processing on the image sequence data to form a preprocessed image set; The feature extraction module 302 is used to perform multi-scale feature extraction on the preprocessed image set through depthwise separable convolution to generate a feature map, perform attention enhancement processing on the feature map to construct a salient region, and perform candidate box filtering based on the salient region to generate defect response parameters. The segment localization module 303 is used to perform spatial pyramid pooling on the feature map to determine fusion features, use the fusion features to perform bounding box regression to locate defect segments, and fuse the defect response parameters based on the defect segment features to form a feature matrix. The map construction module 304 is used to perform pattern clustering on the feature matrix to determine the category distribution, extract weak response regions from the feature matrix to generate early warning markers, generate defect level parameters based on the category distribution and the early warning markers, and extract severity gradient features from the defect level parameters to construct a defect feature map. The detection output module 305 is used to perform feature optimization on the defect response parameters and the feature matrix through multi-scale loss constraints to extract feature sensitivity weights, construct weighted discrimination rules based on the feature sensitivity weights, and integrate the defect feature map through the weighted discrimination rules to output visual detection results.
[0070] The aforementioned machine vision-based workpiece defect detection system 300 can implement a machine vision-based workpiece defect detection method according to the above method embodiments. The options in the above method embodiments are also applicable to this embodiment, and will not be detailed here. The remaining content of this application's embodiments can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.
[0071] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.
Claims
1. A method for detecting defects in a workpiece based on machine vision, the method comprising: include: Collect image sequence data of the workpiece surface, and perform size normalization and data augmentation on the image sequence data to form a preprocessed image set; For the preprocessed image set, multi-scale feature extraction is performed using depthwise separable convolution to generate feature maps. Attention enhancement processing is performed on the feature maps to construct salient regions. Based on the salient regions, candidate boxes are selected to generate defect response parameters. Spatial pyramid pooling is performed on the feature map to determine fusion features. The fusion features are used to perform bounding box regression to locate defect segments. Based on the defect segment features, the defect response parameters are fused to form a feature matrix. Pattern clustering is performed on the feature matrix to determine the category distribution. Weak response regions are extracted from the feature matrix to generate early warning markers. Defect level parameters are generated based on the category distribution and the early warning markers. Severity gradient features are extracted from the defect level parameters to construct a defect feature map. Feature sensitivity weights are extracted by performing feature optimization on the defect response parameters and the feature matrix through multi-scale loss constraints. A weighted discrimination rule is constructed based on the feature sensitivity weights. The defect feature map is then integrated through the weighted discrimination rule to output the visual detection result.
2. The method according to claim 1, characterized in that, The step of generating feature maps by performing multi-scale feature extraction using depthwise separable convolution on the preprocessed image set includes: Channel-wise convolution is performed on the preprocessed image set to generate channel feature responses; The channel compression factor is generated by extracting cross-channel correlation information from the channel feature response. The channel feature response is recalibrated using the channel compression factor to form lightweight features; Feature maps are generated by multi-scale cascade fusion based on the lightweight features.
3. The method according to claim 1, characterized in that, The process of performing attention enhancement processing on the feature map to construct salient regions includes: The feature map is subjected to dual-path separation calculation of channel response and spatial response to generate an initial attention map; Low-response channels are extracted from the initial attention map to generate suppression compensation weights; The initial attention map is weakly enhanced using the suppression compensation weights to form a corrected attention map. High-response regions are selected based on the corrected attention map to form salient regions.
4. The method according to claim 1, characterized in that, The method of using the fused features to perform bounding box regression to locate defective segments includes: A location prediction map is generated by mapping the fused features to the detection space; An initial bounding box is obtained by generating anchor boxes from the location prediction map; Defect morphological features are extracted using the initial bounding box, and aspect ratio constraint regression is performed to form a morphological correction box. Non-maximum suppression is applied to the morphological correction frame to form defect segments.
5. The method according to claim 1, characterized in that, The step of extracting weak response regions from the feature matrix to generate early warning markers includes: Generate a response intensity distribution based on the feature matrix; Identify the response decay trend from the response intensity distribution to locate early defect formation areas; The temporal stability of the early defect formation region is verified to screen for effective weak response regions; Early warning markers are generated based on the effective weak response regions.
6. The method according to claim 1, characterized in that, The step of performing pattern clustering on the feature matrix to determine the category distribution includes: Generate a sample distance metric based on the feature matrix; Extract outlier samples from the sample distance metric to generate boundary correction factors; The center point is initialized using the sample distance metric and combined with the boundary correction factor to form dynamic cluster centers. The samples are assigned to categories based on the dynamic cluster centers to form a class distribution.
7. The method according to claim 1, characterized in that, The step of extracting feature sensitivity weights by performing feature optimization on the defect response parameters and the feature matrix through multi-scale loss constraints includes: The defect response parameters and the feature matrix are used to perform multi-scale decomposition to form hierarchical loss values; The class balance weights are generated by counting the number of samples in each category from the feature matrix. The hierarchical loss values are adaptively weighted using the category balancing weights to form a balanced loss. The feature sensitivity weights are formed by gradient update based on the balance loss.
8. The method according to claim 3, characterized in that, The step of weakly enhancing the initial attention map using the suppression compensation weights to form a corrected attention map includes: Based on the suppression compensation weights, a region difference analysis is performed on the initial attention map to generate an enhancement priority map; The boundary enhancement factor is generated by extracting the weak response region boundary from the enhancement priority map. An enhanced attention map is formed by adaptively stretching the contrast of the initial attention map using the boundary enhancement factor. A corrected attention map is formed by performing response equalization processing based on the enhanced attention map.
9. The method according to claim 5, characterized in that, The step of verifying the temporal stability of the early defect formation region to screen for effective weak response regions includes: Establish multi-frame response correlation based on the aforementioned early defect formation region; Response fluctuation detection is performed along the multi-frame response correlation to form a stability index; The sustained low response region was determined by transient noise interference analysis using the aforementioned stability indices. Based on the persistently low response region, the region validity is calibrated to form an effective weak response region.
10. A workpiece defect detection system based on machine vision, characterized in that, include: The image acquisition module is used to acquire image sequence data of the workpiece surface, and to perform size normalization and data enhancement processing on the image sequence data to form a preprocessed image set; The feature extraction module is used to perform multi-scale feature extraction on the preprocessed image set through depthwise separable convolution to generate feature maps, perform attention enhancement processing on the feature maps to construct salient regions, and perform candidate box filtering based on the salient regions to generate defect response parameters. The segment localization module is used to perform spatial pyramid pooling on the feature map to determine fusion features, use the fusion features to perform bounding box regression to locate defect segments, and fuse the defect response parameters based on the defect segment features to form a feature matrix. The map construction module is used to perform pattern clustering on the feature matrix to determine the category distribution, extract weak response regions from the feature matrix to generate early warning markers, generate defect level parameters based on the category distribution and the early warning markers, and extract severity gradient features from the defect level parameters to construct a defect feature map. The detection output module is used to perform feature optimization on the defect response parameters and the feature matrix through multi-scale loss constraints to extract feature sensitivity weights, construct weighted discrimination rules based on the feature sensitivity weights, and integrate the defect feature map through the weighted discrimination rules to output visual detection results.