A method, system, device, and medium for visual inspection and identification of surface scratches on workpieces.
By acquiring images through multi-angle illumination and polarization cameras, and combining Stokes vector and depth feature fusion networks, the problem of difficulty in identifying weak scratches on highly reflective workpiece surfaces is solved, achieving highly sensitive scratch detection and improving detection robustness and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU POLYTECHNIC
- Filing Date
- 2026-04-23
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies are insufficient to effectively identify subtle scratches on highly reflective workpiece surfaces. Traditional methods lack robustness under single illumination angles and have high system complexity, making it impossible to achieve high-sensitivity detection without increasing system complexity.
Multiple original images with different polarization directions were acquired using a multi-angle illumination light source array and a polarization camera. The total intensity, degree of polarization, and polarization angle of the images were extracted using Stokes vectors. A deep feature fusion network was constructed, and the scratch edge response was enhanced using a convolutional neural network and spatial attention mechanism to perform semantic segmentation and identify weak scratches.
It achieves highly sensitive detection of subtle scratches on highly reflective workpiece surfaces without increasing system complexity, improves the comprehensive capture capability of subtle scratches with different orientations, enhances the contrast between defects and background, and significantly improves detection sensitivity and accuracy.
Smart Images

Figure CN122090460A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual inspection technology, and in particular to a method, system, device and medium for visual inspection of workpiece surface scratch recognition. Background Technology
[0002] In the field of high-end equipment manufacturing, such as the surfaces of workpieces like aero-engine blades, precision molds, and automotive paint, the presence of minor scratches can directly affect product performance and service life. Because these workpieces usually have high reflectivity, the optical contrast between the microscopic scratches on the surface and the background material is extremely low, making it difficult for traditional machine vision inspection methods based on grayscale images to effectively identify them.
[0003] In existing technologies, high-angle ring light sources or multispectral imaging are commonly used to enhance defect contrast. However, under a single illumination angle, the visibility of scratches is highly dependent on the geometric relationship between their orientation and the direction of the light source. Furthermore, specular reflection can easily lead to local overexposure of the image, completely obscuring weak scratch signals. While methods such as speckle interferometry offer high detection accuracy, they are complex and costly. Polarization imaging can enhance material contrast, but it provides incomplete information under a single illumination angle and lacks an effective information fusion mechanism, resulting in insufficient detection robustness. Therefore, achieving high-sensitivity detection of weak scratches on highly reflective workpiece surfaces without increasing system complexity has become a challenge for the industry. Summary of the Invention
[0004] Based on this, this application provides a visual inspection method, system, device, and medium for identifying surface scratches on highly reflective workpieces with high sensitivity.
[0005] In a first aspect, this application provides a method for visually inspecting and identifying scratches on the surface of a workpiece, comprising the following steps: Control the multi-angle illumination source array to sequentially illuminate the surface of the workpiece under different incident angles, and simultaneously control the polarization camera to acquire multiple original images with different polarization directions at each illumination angle; Based on all the original images, the Stokes vectors under different illumination angles are determined, and the total intensity image, polarization degree image, and polarization angle image are extracted from each Stokes vector; The total intensity image, polarization degree image, and polarization angle image corresponding to the same illumination angle are normalized respectively to make the value range consistent, and then the channels are superimposed to obtain the multimodal feature image under the illumination angle. A deep feature fusion network is constructed. Multimodal feature images under various lighting angles are used as input. A convolutional neural network is used to extract spatial feature maps for each lighting angle. The response of scratch edges in each spatial feature map is enhanced through a spatial attention mechanism. Then, all enhanced spatial feature maps are adaptively weighted and fused to generate a fused feature map. Semantic segmentation is performed on the fused feature map to identify and mark the weak scratch areas on the surface of the detected workpiece.
[0006] In some embodiments, determining the Stokes vector at different illumination angles based on all the original images specifically includes: Based on multiple original images with different polarization directions at each illumination angle, calculate the Stokes vector component image at each illumination angle; The calculated Stokes vector component image is output as the Stokes vector at the corresponding illumination angle.
[0007] In some embodiments, extracting the total intensity image, polarization degree image, and polarization angle image from the various Stokes vectors specifically includes: Extract the total intensity image at each illumination angle from the Stokes vector at each illumination angle; Calculate the degree of polarization image and the angle of polarization image at each illumination angle based on the Stokes vector at each illumination angle.
[0008] In some embodiments, the total intensity image, polarization degree image, and polarization angle image corresponding to the same illumination angle are normalized respectively to ensure consistent value ranges before channel superposition to obtain a multimodal feature image at that illumination angle. Specifically, this includes: The total intensity image, polarization degree image, and polarization angle image under the same illumination angle are normalized to make the pixel value range of the three consistent. The normalized total intensity image, normalized polarization degree image, and normalized polarization angle image are overlaid to generate a multimodal feature image at this illumination angle.
[0009] In some embodiments, the extraction of spatial feature maps for each illumination angle using a convolutional neural network specifically includes: The multimodal feature images under each illumination angle are respectively input into a convolutional neural network with shared weights; The spatial feature map corresponding to each illumination angle is obtained by forward propagation calculation through the convolutional neural network.
[0010] In some embodiments, enhancing the response of scratch edges in each spatial feature map through a spatial attention mechanism specifically includes: Generate a corresponding spatial attention weight map for the spatial feature map of each illumination angle; Each spatial attention weight map is multiplied element-wise with its corresponding spatial feature map to enhance the response of scratch edges in each spatial feature map.
[0011] In some embodiments, adaptive weighted fusion of all enhanced spatial feature maps to generate a fused feature map specifically includes: For each enhanced spatial feature map at each illumination angle, calculate the global descriptive vector along the channel dimension. Adaptive fusion weights for each lighting angle are learned based on each global description vector; The spatial feature maps enhanced by all lighting angles are weighted and summed according to the learned adaptive fusion weights to generate the final fused feature map.
[0012] Secondly, this application provides a visual inspection system for identifying surface scratches on a workpiece, comprising: The acquisition module is used to control the multi-angle illumination source array to sequentially illuminate the surface of the workpiece under different incident angles, and simultaneously control the polarization camera to acquire multiple original images with different polarization directions at each illumination angle; The processing module is used to determine the Stokes vectors under different illumination angles based on all the original images, and extract the total intensity image, polarization degree image and polarization angle image from each Stokes vector; The processing module is also used to normalize the total intensity image, polarization degree image and polarization angle image corresponding to the same illumination angle, and then perform channel superposition after making the value range consistent to obtain the multimodal feature image under the illumination angle. The processing module is also used to construct a deep feature fusion network. Taking the multimodal feature images under each illumination angle as input, a convolutional neural network is used to extract the spatial feature map of each illumination angle. The response of the scratch edge in each spatial feature map is enhanced by a spatial attention mechanism. Then, all enhanced spatial feature maps are adaptively weighted and fused to generate a fused feature map. The execution module is used to perform semantic segmentation on the fused feature map, identify and mark the weak scratch areas on the surface of the detected workpiece.
[0013] Thirdly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described visual inspection method for identifying scratches on the surface of a workpiece.
[0014] Fourthly, this application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described visual inspection method for identifying scratches on the surface of a workpiece.
[0015] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: The visual inspection method, system, equipment, and medium for identifying surface scratches on workpieces provided in this application firstly control a multi-angle illumination light source array to sequentially illuminate the workpiece surface at different incident angles, and simultaneously control a polarization camera to acquire multiple original images with different polarization directions at each illumination angle. This step enables simultaneous acquisition of multi-angle illumination and polarization images, obtaining original images with multiple polarization directions at different incident angles, thereby improving the comprehensive capture capability of subtle scratches with different orientations. Secondly, based on all the original images, the Stokes vectors at different illumination angles are determined, and the total intensity image and polarization vectors are extracted from each Stokes vector. The vibration and polarization angle images are extracted. This step calculates the Stokes vector from the original image and extracts three types of feature images: total intensity, degree of polarization, and polarization angle. It fully utilizes the sensitivity of polarization information to material changes, effectively separating subtle scratches from a highly reflective background and enhancing the contrast between defects and the background. Subsequently, the total intensity, degree of polarization, and polarization angle images corresponding to the same illumination angle are normalized to ensure consistent value ranges before channel overlay, resulting in a multimodal feature image for that illumination angle. This step normalizes and overlays the three types of feature images at the same angle, eliminating discrepancies in physical quantities. By differentiating dimensions and fusing multi-dimensional information, standardized multimodal feature images are formed, thus providing high-quality input data for deep learning networks. Then, a deep feature fusion network is constructed, using multimodal feature images from various illumination angles as input. A convolutional neural network extracts spatial feature maps for each illumination angle, and a spatial attention mechanism enhances the response of scratch edges in each spatial feature map. Finally, all enhanced spatial feature maps are adaptively weighted and fused to generate a fused feature map. This step enables adaptive extraction and fusion of multi-angle spatial features, strengthens scratch edge responses through spatial attention, and combines channel attention... By dynamically weighting the contributions from each angle, a fused feature map highlighting subtle defects is generated, significantly improving the network's sensitivity to low-contrast scratches. Finally, semantic segmentation is performed on the fused feature map to identify and label the subtle scratch areas on the workpiece surface. This step enables pixel-level classification and post-processing of the fused feature map, accurately segmenting the subtle scratch areas and labeling their geometric parameters, thus visualizing the detection results intuitively. This completes the entire process from the original image to defect recognition. In summary, the solution proposed in this application can achieve high-sensitivity detection of subtle scratches on highly reflective workpiece surfaces without increasing system complexity. Attached Figure Description
[0016] Figure 1 This is an exemplary flowchart of a visual inspection method for identifying surface scratches on a workpiece, according to some embodiments of this application. Figure 2 This is a schematic diagram illustrating an application scenario of a surface scratch recognition data processing system according to some embodiments of this application; Figure 3 This is a schematic flowchart illustrating the generation of fused feature maps according to some embodiments of this application; Figure 4 This is a schematic diagram of the structure of a visual inspection workpiece surface scratch recognition system according to some embodiments of this application; Figure 5 This is a schematic diagram of the structure of a computer device for implementing a visual inspection method for identifying scratches on the surface of a workpiece, according to some embodiments of this application. Detailed Implementation
[0017] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0018] refer to Figure 1 The figure is an exemplary flowchart of a visual inspection method for identifying scratches on a workpiece surface according to some embodiments of this application. The visual inspection method for identifying scratches on a workpiece surface mainly includes the following steps: In step 101, the multi-angle illumination light source array is controlled to sequentially illuminate the surface of the workpiece at different incident angles, and the polarization camera is simultaneously controlled to acquire multiple original images with different polarization directions at each illumination angle.
[0019] In practical implementation, controlling a multi-angle illumination array to sequentially illuminate the surface of the workpiece at different incident angles, while simultaneously controlling a polarization camera to acquire multiple raw images with different polarization directions at each illumination angle, can be achieved as follows: First, a ring-shaped LED light source array is installed above the inspection station. This array contains multiple independently controlled illumination units, each distributed at different zenith angles centered on the workpiece. Specifically, these are set at three incident angles: 30°, 45°, and 60°. Each incident angle corresponds to a set of ring-shaped LED beads. All light sources use unpolarized white light to avoid interference. Additional polarization modulation is introduced; simultaneously, a focal plane polarization camera is vertically mounted directly above the workpiece. This camera's sensor pixels integrate a micro-polarizer array, enabling simultaneous acquisition of images from four polarization directions (0°, 45°, 90°, and 135°) in a single exposure. The camera resolution is selected as 5 megapixels based on the required detection accuracy, and a telecentric lens is used to ensure a uniform imaging field of view and distortion control within 0.1%. Next, a programmable logic controller (PLC) or embedded microcontroller sequentially illuminates the light sources at each angle according to a preset timing sequence. Specifically, the timing sequence is as follows: first, the 30° incident angle light source is illuminated and stabilized for 10 milliseconds. The controller simultaneously sends a trigger signal to the camera, which immediately acquires one frame and outputs four original grayscale images corresponding to polarization directions of 0°, 45°, 90°, and 135°, respectively. Each image is saved in 8-bit BMP format and stored in the local cache according to the naming rule of "angle_polarization direction". Then, the 30° light source is turned off, and the 45° light source is turned on after a 5-millisecond interval, repeating the above trigger acquisition and storage process. Finally, the 60° light source is turned on to complete the third acquisition. Through the above control flow, a set of four original images with different polarization directions corresponding to each illumination angle is finally obtained, namely 30°, 45°, 90°, and 135° polarization direction, respectively. Images at incident angles of I30°_0°, I30°_45°, I30°_90°, and I30°_135°, images at an incident angle of 45° of I45°_0°, I45°_45°, I45°_90°, and I45°_135°, and images at an incident angle of 60° of I60°_0°, I60°_45°, I60°_90°, and I60°_135° are spatially fully registered and have one-to-one pixel correspondence, providing an accurate data foundation for subsequent polarization parameter calculations. Other methods can also be used in other embodiments, which are not limited here.
[0020] It should be noted that the above steps can achieve simultaneous acquisition of multi-angle illumination and polarization images, obtaining original images of multiple polarization directions under different incident angles, thereby improving the comprehensive capture capability of weak scratches with different orientations.
[0021] In some embodiments, reference Figure 2As shown in the figure, this figure is a schematic diagram of the application scenario of the surface scratch recognition data processing system shown in some embodiments of this application. The figure includes three main components: acquisition device, server and data storage device. The acquisition device is responsible for collecting multiple original images with different polarization directions and sending the acquired original images with different polarization directions to the server through a communication network. The surface scratch recognition data processing system runs in the server. The server stores the processing results in the data storage device and visualizes them.
[0022] In step 102, Stokes vectors under different illumination angles are determined based on all the original images, and the total intensity image, polarization degree image and polarization angle image are extracted from each Stokes vector.
[0023] In some embodiments, determining the Stokes vector at different illumination angles based on all the original images can be achieved by the following steps: Based on multiple original images with different polarization directions at each illumination angle, calculate the Stokes vector component image at each illumination angle; The calculated Stokes vector component image is output as the Stokes vector at the corresponding illumination angle.
[0024] It should be noted that the Stokes vector in this application is a physical quantity used to comprehensively describe the polarization state of the beam. In this method, it is composed of three component images: S0, S1, and S2. The Stokes vector component images are the single-channel images that constitute the Stokes vector, corresponding to the S0 component (total light intensity), the S1 component (horizontal and vertical polarization difference), and the S2 component (45° and 135° polarization difference), respectively, and are used to store polarization information in different dimensions.
[0025] In specific implementation, the Stokes vector component image for each illumination angle can be calculated based on multiple original images with different polarization directions at each illumination angle. This can be achieved as follows: For each illumination angle, four original grayscale images corresponding to the 0°, 45°, 90°, and 135° polarization directions are read. For example, for a 30° incident angle, four images are read: I30°_0°, I30°_45°, I30°_90°, and I30°_135°. Then, the grayscale values of the four images are weighted and combined at the pixel level. Specifically, for each pixel position, the S1 component value of the pixel is obtained by subtracting the grayscale value of the 0° and 90° polarization directions, and the S1 component value of the pixel is obtained by subtracting the grayscale value of the 45° and 135° polarization directions. The S2 component value of the pixel is obtained, and then the sum of the gray values in the four directions of 0°, 45°, 90°, and 135° is divided by 2 to obtain the S0 component value of the pixel. After performing the above calculation on all pixels of the entire image, three single-channel images with the same size as the original image are generated, corresponding to the S0 component image, S1 component image, and S2 component image of the Stokes vector, respectively. The S0 component image reflects the total light intensity, the S1 component image reflects the polarization difference in the horizontal and vertical directions, and the S2 component image reflects the polarization difference in the 45° and 135° directions. The above process is repeated for all illumination angles to obtain the S0 component image, S1 component image, and S2 component image corresponding to each illumination angle. Other methods can also be used in other embodiments, which are not limited here.
[0026] In specific implementation, the calculated Stokes vector component images can be output as Stokes vectors under the corresponding illumination angles in the following way: the S0, S1, and S2 component images calculated for each illumination angle are stacked in channel order to form a three-channel image, and this three-channel image is output as the Stokes vector under that illumination angle to a temporary storage area in memory. Angle labels, such as 30°, 45°, and 60°, can be added to the three-channel images, or the three single-channel images can be stored as independent files in a specified directory, and an index table can be established in memory with the illumination angle as the key and the file paths of the S0, S1, and S2 images as the values. This allows for quick indexing of the corresponding S0, S1, and S2 component images according to the illumination angle during subsequent processing. During output, it is ensured that the Stokes vector for each illumination angle strictly corresponds to its acquisition angle label, providing complete and traceable data input for subsequent extraction of the total intensity image, polarization degree image, and polarization angle image from the Stokes vector. Other methods can also be used in other embodiments, and are not limited here.
[0027] In some embodiments, extracting the total intensity image, polarization degree image, and polarization angle image from the individual Stokes vectors can be achieved by the following steps: Extract the total intensity image at each illumination angle from the Stokes vector at each illumination angle; Calculate the degree of polarization image and the angle of polarization image at each illumination angle based on the Stokes vector at each illumination angle.
[0028] It should be noted that the total intensity image in this application, i.e., the S0 component image, reflects the total intensity of light reflected from the workpiece surface, which is equivalent to a traditional grayscale image and provides basic brightness information for multimodal features; the polarization degree image represents the degree of polarization of light at each pixel position, calculated from S0, S1, and S2, and is sensitive to material changes, used to enhance the contrast between subtle scratches and the background; the polarization angle image represents the polarization direction angle of light at each pixel position, calculated from S1 and S2, and provides directional features for scratch detection.
[0029] In practice, extracting the total intensity image for each illumination angle from the Stokes vector can be achieved as follows: Read the stored Stokes vector for each illumination angle. This Stokes vector is either a three-channel image containing S0, S1, and S2 component images, or three independent single-channel images. For each illumination angle, directly extract the S0 component image from its Stokes vector as the total intensity image for that illumination angle. Specifically, if the Stokes vector is stored as a three-channel image, extract the image data of the first channel through an image channel separation operation. This refers to the S0 component image. If the Stokes vector is stored as three independent single-channel images (S0, S1, and S2), the S0 component image file is read directly. The extracted S0 component images are named in the format of total intensity image_angle, such as total intensity image_30°, total intensity image_45°, and total intensity image_60°, and stored in memory or a temporary folder. At the same time, a lookup table indexed by the illumination angle is established to record the storage location of each total intensity image, ensuring that the total intensity image of each illumination angle strictly corresponds to its acquisition angle label. Other methods can also be used in other embodiments, which are not limited here.
[0030] In specific implementation, the polarization degree image and polarization angle image corresponding to each illumination angle can be calculated based on the Stokes vector at that illumination angle in the following way: For each illumination angle, firstly, read the S0 component image, S1 component image, and S2 component image from the Stokes vector of that illumination angle. Then, perform traversal calculations at the pixel level for each pixel position: For the polarization degree image, calculate the sum of the squares of the S1 value and the S2 value of the pixel, take the square root of the sum, and divide it by the S0 value of the pixel. The result is the polarization degree value of the pixel. During the calculation process, division by zero protection is required. When the S0 value is less than a preset threshold, such as 0.001, the polarization degree value of the pixel is directly set to 0. For the polarization angle image, calculate the ratio of the S2 value to the S1 value of the pixel, perform an arctangent operation on the ratio, and divide it by 2. The result is expressed in radians, and its value is... The range is from -π / 2 to π / 2. To facilitate subsequent image processing, the radian value is converted to an angle value. Specifically, the radian value is multiplied by 180 and then divided by π to obtain an angle value within the range of 0 to 180 degrees. After performing the above calculation on all pixels of the entire image, two single-channel images with the same size as the original image are generated, which are respectively used as the polarization degree image and polarization angle image under the illumination angle. The polarization degree image is named polarization degree image_angle format, and the polarization angle image is named polarization angle image_angle format, for example, polarization degree image_30° and polarization angle image_30°. They are stored in the same directory as the total intensity image under the illumination angle or a unified index relationship is established to ensure that there is a one-to-one correspondence between the pixels of the total intensity image, polarization degree image, and polarization angle image under each illumination angle. Other methods can also be used in other embodiments, which are not limited here.
[0031] It should be noted that the above steps can solve the Stokes vector from the original image and extract three types of feature images: total intensity, degree of polarization, and polarization angle. By making full use of the sensitivity of polarization information to material changes, the faint scratches can be effectively separated from the highly reflective background, thereby enhancing the contrast between the defect and the background.
[0032] In step 103, the total intensity image, polarization degree image and polarization angle image corresponding to the same illumination angle are normalized respectively to make the value range consistent, and then the channels are superimposed to obtain the multimodal feature image under the illumination angle.
[0033] In some embodiments, the total intensity image, polarization degree image, and polarization angle image corresponding to the same illumination angle are normalized to ensure consistent value ranges before channel superposition to obtain a multimodal feature image at that illumination angle. This can be achieved by the following steps: The total intensity image, polarization degree image, and polarization angle image under the same illumination angle are normalized to make the pixel value range of the three consistent. The normalized total intensity image, normalized polarization degree image, and normalized polarization angle image are overlaid to generate a multimodal feature image at this illumination angle.
[0034] It should be noted that the multimodal feature image in this application is a three-channel image formed by superimposing the total intensity image, polarization degree image and polarization angle image under the same illumination angle, which is used to fuse multiple physical features as input to a deep learning network.
[0035] In practice, normalizing the total intensity image, polarization degree image, and polarization angle image under the same illumination angle to ensure consistent pixel value ranges can be achieved as follows: First, read the total intensity image, polarization degree image, and polarization angle image under the same illumination angle. For example, for a 30° incident angle, read the total intensity image _30°, polarization degree image _30°, and polarization angle image _30°. Then, for the total intensity image, if its pixel data type is an 8-bit unsigned integer with a value range of 0 to 255, divide each pixel value by 255 to convert it to a floating-point number in the range of 0 to 1. If the total intensity image is already a floating-point number, use it directly. For the polarization degree image, its pixel values are already in the range of 0 to 1, so it only needs to be converted to the same floating-point data type as the total intensity image. Any pixel type is acceptable; for polarization angle images, if they are stored in degrees and the value range is 0 to 180, then divide each pixel value by 180 to map it to the 0 to 1 interval. If they are stored in radians and the value range is -2 / π to 2 / π, then first add 2 / π to each pixel value to make its value range 0 to π, and then divide it by π to map it to the 0 to 1 interval. After completing the above normalization operation, three floating-point images with pixel value ranges of 0 to 1 and consistent data types are obtained. These are used as the normalized total intensity image, the normalized polarization degree image, and the normalized polarization angle image, respectively. They are temporarily stored in memory or stored in association with the original image to ensure that the three are completely consistent in spatial resolution and that the pixel positions correspond one-to-one. Other methods can also be used in other embodiments, which are not limited here.
[0036] In specific implementation, the normalized total intensity image, normalized polarization degree image, and normalized polarization angle image are superimposed to generate a multimodal feature image for that illumination angle. This can be achieved as follows: For each illumination angle, a three-channel empty image matrix is created. The height and width of this empty image matrix are the same as the normalized single-channel image, and the data type is floating-point. The normalized total intensity image is assigned as the data of the first channel to the first channel of this matrix, the normalized polarization degree image is assigned as the data of the second channel to the second channel, and the normalized polarization angle image is superimposed to the third channel. Data from the third channel is assigned to the third channel to form a three-channel multimodal feature image. This multimodal feature image is named "multimodal feature image_angle format", for example, "multimodal feature image_30°", and stored in memory or a temporary file. At the same time, a lookup table indexed by the illumination angle is created to record the storage location of each multimodal feature image. The above operation is repeated for all illumination angles to obtain a set of multimodal feature images. Each image integrates intensity information, polarization degree information, and polarization angle information at the same angle. Other methods can also be used in other embodiments, which are not limited here.
[0037] It should be noted that the above steps can normalize and overlay the three types of feature images at the same angle, eliminate the dimensional differences of different physical quantities and fuse multi-dimensional information to form a standardized multimodal feature image, thereby providing high-quality input data for deep learning networks.
[0038] In step 104, a deep feature fusion network is constructed. The multimodal feature images under each illumination angle are used as input. A convolutional neural network is used to extract the spatial feature map of each illumination angle. The response of the scratch edge in each spatial feature map is enhanced by a spatial attention mechanism. Then, all enhanced spatial feature maps are adaptively weighted and fused to generate a fused feature map.
[0039] In practical implementation, a deep feature fusion network is constructed, using multimodal feature images from various illumination angles as input. This can be achieved as follows: First, a multi-input network structure is defined based on a deep learning framework such as PyTorch or TensorFlow. This network structure contains multiple parallel input branches, the number of which is the same as the preset number of illumination angles. For example, three illumination angles correspond to three input branches. Each input branch receives a multimodal feature image from the corresponding illumination angle. This multimodal feature image is in three-channel floating-point format, with a size of H×W×3, where H and W are the image height and width, respectively. Then, a convolutional neural network with shared weights is configured for all input branches as the feature extraction backbone. This feature extraction backbone uses a ResNet-18 model pre-trained on the ImageNet dataset, and removes... Its global average pooling layer and fully connected layer retain only the first four stages of convolutional layers, and the output feature map size is H / 8×W / 8 with 512 channels. In the specific implementation, the multimodal feature images received by each branch are forward-propagated through this shared backbone network to obtain spatial feature maps corresponding to each illumination angle, denoted as F_θ1, F_θ2, and F_θ3. These feature maps have the same size and number of channels and share the same network parameters. At the same time, in the code implementation, parameter consistency is ensured by applying a copy of the shared backbone network to each input branch or by using weight copying. After the network is built, the multimodal feature images under each illumination angle are passed into the corresponding branches in angular order to obtain the spatial feature map of each angle for subsequent module processing. Other methods can also be used in other embodiments, which are not limited here.
[0040] In some embodiments, the extraction of spatial feature maps for each illumination angle using a convolutional neural network can be achieved through the following steps: The multimodal feature images under each illumination angle are respectively input into a convolutional neural network with shared weights; The spatial feature map corresponding to each illumination angle is obtained by forward propagation calculation through the convolutional neural network.
[0041] It should be noted that the spatial feature map in this application is a high-dimensional feature representation extracted from multimodal feature images by a convolutional neural network, which encodes the spatial structure information of the image.
[0042] In practice, the multimodal feature images at each illumination angle are input into a shared-weight convolutional neural network as follows: First, multimodal feature images at all illumination angles are read. For example, for angles of 30°, 45°, and 60°, multimodal feature images _30°, _45°, and _60° are read respectively. Each image is in three-channel floating-point format, with dimensions denoted as H×W×3. Then, a shared-weight convolutional neural network is defined in the deep learning framework. This network uses a ResNet-18 model pre-trained on the ImageNet dataset as the feature extraction backbone, and its global average is removed. Pooling layers and fully connected layers are used, with only the first four stages of convolutional layers retained. The output feature map has a size of H / 8×W / 8 and 512 channels. Next, parallel input branches with the same number of illumination angles are created in the code implementation. Each branch receives a multimodal feature image, and the input data of all branches are organized into a batch tensor with dimensions of N×3×H×W, where N is the number of illumination angles, so that it can be forward propagated through the shared network in one go. Finally, the organized batch tensor is input into the convolutional neural network with shared weights, so that the forward propagation calculation can be performed in parallel for each illumination angle. Other methods can also be used in other embodiments, which are not limited here.
[0043] In specific implementation, the spatial feature map corresponding to each illumination angle can be obtained through the forward propagation calculation of the convolutional neural network in the following way: After passing the multimodal feature images of each illumination angle as batch input to the convolutional neural network with shared weights, the convolutional neural network automatically performs layer-by-layer convolution, batch normalization, nonlinear activation and downsampling operations, and finally outputs a batch-processed feature map tensor with dimensions of N×512×H / 8×W / 8, where N is the number of illumination angles, 512 is the number of channels, and H / 8 and W / 8 are the spatial dimensions; then, the batch tensor is split into N independent feature maps along the first dimension, each feature map corresponding to an illumination angle, denoted as F_30°, F_45°, and F_60° respectively. These feature maps are the spatial feature maps corresponding to each illumination angle. They have the same size and number of channels and share the same convolution kernel parameters. Other methods can also be used in other embodiments, which are not limited here.
[0044] In some embodiments, enhancing the response of scratch edges in each spatial feature map through a spatial attention mechanism can be achieved by the following steps: Generate a corresponding spatial attention weight map for the spatial feature map of each illumination angle; Each spatial attention weight map is multiplied element-wise with its corresponding spatial feature map to enhance the response of scratch edges in each spatial feature map.
[0045] It should be noted that the spatial attention weight map in this application is a weight map of the same size as the spatial feature map generated by the spatial attention mechanism. Each pixel value represents the importance of that location and is used to enhance the scratch edge response and suppress background noise.
[0046] In specific implementation, generating corresponding spatial attention weight maps for the spatial feature maps of each illumination angle can be achieved as follows: For the spatial feature maps of each illumination angle, firstly, global max pooling and global average pooling are performed on the spatial feature maps in the channel dimension to obtain two two-dimensional feature maps of size H / 8×W / 8×1. Then, these two feature maps are concatenated in the channel dimension to obtain a concatenated feature map of size H / 8×W / 8×2. Next, this concatenated feature map is input into a 7×7 convolutional layer, which outputs 1 channel and maps the output value through a Sigmoid activation function. The process is repeated from 0 to 1, resulting in a spatial attention weight map with dimensions H / 8×W / 8×1. The value of each pixel in this weight map represents the importance of the corresponding position. The scratch edge region is assigned a higher weight value because of the large difference in response between global max pooling and global average pooling. The above operation is repeated for the spatial feature maps of all illumination angles to obtain the spatial attention weight map corresponding to each illumination angle, denoted as A_30°, A_45°, and A_60°. These spatial attention weight maps have the same spatial size as the corresponding spatial feature maps. Other methods can also be used in other embodiments, which are not limited here.
[0047] In specific implementation, the element-wise multiplication of each spatial attention weight map with the corresponding spatial feature map to enhance the response of scratch edges in each spatial feature map can be achieved in the following way: For each illumination angle, its spatial attention weight map is multiplied with the corresponding spatial feature map element-wise. Specifically, the spatial attention weight map is first copied in the channel dimension to expand its channel number to the same 512 channels as the spatial feature map. Then, a pixel-wise multiplication operation is performed, that is, each pixel value in each channel is multiplied by the attention weight value of the corresponding spatial position to obtain the enhanced spatial feature map. This operation makes the scratch edge region in the spatial feature map, which originally had a weak response, amplified due to the higher attention weight, while the background region is suppressed due to the lower weight. After performing the above element-wise multiplication for all illumination angles, the enhanced spatial feature maps for each illumination angle are obtained, denoted as F'_30°, F'_45°, and F'_60°. The size and number of channels of these feature maps are consistent with the original spatial feature map, but the response of scratch edges is significantly enhanced. Other methods can also be used in other embodiments, which are not limited here.
[0048] In some embodiments, reference Figure 3As shown, this figure is a schematic diagram of the process for generating fused feature maps in some embodiments of this application. In this embodiment, adaptive weighted fusion is performed on all enhanced spatial feature maps to generate the fused feature map. The following steps can be used to achieve this: In step 1031, a global description vector for the channel dimension is calculated for each spatial feature map enhanced at each illumination angle. In step 1032, adaptive fusion weights for each illumination angle are learned based on each global description vector; In step 1033, the spatial feature maps enhanced by all lighting angles are weighted and summed according to the learned adaptive fusion weights to generate the final fusion feature map.
[0049] It should be noted that the global description vector in this application is a vector obtained by global pooling of the enhanced spatial feature map, representing the overall response intensity of each channel, and is used to learn the adaptive fusion weights for each illumination angle; the adaptive fusion weights are a set of scalar weights learned by the global description vector through a fully connected layer, with one weight corresponding to each illumination angle, used to dynamically adjust the contribution ratio of each angle feature during fusion; the fusion feature map is a feature map obtained by weighting and summing the enhanced spatial feature maps of each illumination angle according to the adaptive fusion weights, which integrates complementary information from multiple angles and highlights weak scratch features.
[0050] In specific implementation, the global description vector of the channel dimension for each enhanced spatial feature map at each illumination angle can be calculated in the following way: For each enhanced spatial feature map at each illumination angle, a global average pooling operation is first performed on it, that is, the average value of each channel is calculated in the spatial dimensions (height and width) to obtain the channel description vector of the corresponding spatial feature map, which has a dimension of 512×1×1. Each element in this channel description vector represents the global response intensity of the corresponding channel. At the same time, in order to retain richer statistical information, a global max pooling operation can also be performed in parallel to obtain another 512×1×1 channel description vector. Then, the two description vectors are added or concatenated. This embodiment uses the addition method to fuse the two to enhance the description capability. The above operation is repeated for all illumination angles to obtain the global description vector corresponding to each illumination angle, denoted as V_30°, V_45°, and V_60°. These vectors capture the global characteristics of the feature map at each angle in the channel dimension. Other methods can also be used in other embodiments, which are not limited here.
[0051] In specific implementation, the adaptive fusion weights for each illumination angle can be learned from the global description vectors as follows: All global description vectors for illumination angles are concatenated along the channel dimension to form a combined vector of dimension (512×N)×1×1, where N is the number of illumination angles. This combined vector is then input into a lightweight network consisting of two fully connected layers. The first fully connected layer compresses the dimension to 512 / N (e.g., 512 / 3≈170) and passes it through the ReLU activation function. The second fully connected layer restores the dimension to N and passes it through the Softmax activation function, ultimately outputting an N-dimensional weight vector. Each element in this weight vector corresponds to an adaptive fusion weight for an illumination angle, and the sum of all weights is 1. This weight learning process is end-to-end trained, allowing the network to automatically adjust the contribution of each angle based on the input features, giving higher weights to angles more favorable for scratch detection. Other methods can also be used in other embodiments, which are not limited here.
[0052] In specific implementation, the final fused feature map is generated by weighted summation of the spatial feature maps enhanced by all illumination angles according to the learned adaptive fusion weights. This can be achieved as follows: First, the learned N-dimensional adaptive fusion weight vector is decomposed into scalar weights for each illumination angle, such as w_30°, w_45°, and w_60°. Then, for each illumination angle, the enhanced spatial feature map is multiplied by the corresponding scalar weight for each channel to obtain a weighted feature map. Finally, all weighted feature maps are summed at the element level, i.e., pixel values at the same spatial location and in the same channel are directly summed to generate the final fused feature map F_fusion, which has the same size as a single enhanced spatial feature map, i.e., H / 8 × W / 8, with 512 channels. This fused feature map integrates complementary information from multi-angle illumination and strengthens the contribution of key angles through adaptive weights. Other methods can also be used in other embodiments, which are not limited here.
[0053] It should be noted that the above steps can achieve adaptive extraction and fusion of multi-angle spatial features, enhance the scratch edge response through spatial attention mechanism, and combine channel attention to dynamically weight the contribution of each angle to generate a fusion feature map that highlights weak defects, thereby significantly improving the network's detection sensitivity for low-contrast scratches.
[0054] In step 105, semantic segmentation is performed on the fused feature map to identify and mark the weak scratch areas on the surface of the workpiece.
[0055] In some embodiments, semantic segmentation of the fused feature map to identify and label the weak scratch regions on the surface of the workpiece can be achieved by the following steps: A semantic segmentation decoder is constructed, and then the fused feature map is upsampled and classified at the pixel level to obtain a scratch probability map; The scratch probability map is post-processed to identify and mark the weak scratch areas on the surface of the workpiece.
[0056] In specific implementation, a semantic segmentation decoder is constructed, and then the fused feature map is upsampled and classified at the pixel level to obtain the scratch probability map. This can be achieved in the following way: First, a semantic segmentation decoder with a feature pyramid network structure is defined in the deep learning framework. The semantic segmentation decoder takes the fused feature map F_fusion as input, with a size of H / 8×W / 8 and 512 channels. The semantic segmentation decoder first compresses the number of channels of the fused feature map to 256 through a 1×1 convolutional layer to reduce the computational load, and then performs a 2x upsampling using bilinear interpolation to obtain a feature map with a size of H / 4×W / 4. At the same time, a low-level feature map is extracted from the corresponding layer of the shared convolutional neural network feature extraction backbone used in step 104. For example, a low-level feature map with a size of H / 4×W / 4 and 128 channels is extracted from the second-stage output of ResNet-18. After adjusting the number of channels of this low-level feature map to 256 through a 1×1 convolutional layer, it is fused element-wise with the upsampled feature map. Then, a 2x upsampling is performed again to obtain a feature map with a size of H / 4×W / 4. The feature map is H / 2×W / 2, and similarly fused with a low-level feature map of size H / 2×W / 2 with 64 channels extracted from the first stage of ResNet-18. Finally, a 2x upsampling is performed to restore the original image size H×W, and a 1×1 convolutional layer is used to map the feature map into a 2-channel output, corresponding to the scratch category and the background category, respectively. During the model training phase, the 2-channel output is followed by a Softmax activation function to obtain the probability value of each pixel belonging to scratch and background, and the cross-entropy loss and Dice loss are calculated with the labeled real scratch mask for end-to-end training. During the model inference phase, the trained decoder is applied to the fused feature map F_fusion. After the above upsampling and convolution operations, the probability value of the scratch category in the Softmax output is taken as the scratch probability of each pixel, and finally a single-channel scratch probability map with the same size as the original image is generated. In this map, each pixel value is between 0 and 1, indicating the possibility of a scratch at that location. Other methods can also be used in other embodiments, which are not limited here.
[0057] In specific implementation, post-processing the scratch probability map to identify and mark the weak scratch areas on the workpiece surface can be achieved in the following way: First, threshold segmentation is performed on the scratch probability map. A fixed threshold, such as 0.5, is set, and pixels with a probability value greater than or equal to 0.5 are marked as 1 to represent scratches, and pixels with a probability value less than 0.5 are marked as 0 to represent background, resulting in a binary scratch mask image. Then, morphological opening is performed on the binary scratch mask image, using a 3×3 rectangular structuring element for erosion followed by dilation to eliminate isolated small points and tiny erroneous connections caused by noise. Finally, connected component segmentation is performed on the morphologically processed binary scratch mask image. The analysis employs an eight-neighbor connectivity criterion to mark all independent connected regions, and calculates the area (i.e., the number of pixels) of each connected region. An area threshold, such as 50 pixels, is set, and connected regions with areas smaller than this threshold are identified as noise and filtered out. The remaining connected regions after filtering are retained as the final weak scratch regions. Finally, the contours of the retained weak scratch regions are extracted, or drawn on the original workpiece image in the form of a minimum bounding rectangle. Simultaneously, the geometric parameters such as the center coordinates, area, and direction of each scratch region are output, completing the identification and marking of weak scratches on the workpiece surface. Other methods can also be used in other embodiments, which are not limited here.
[0058] It should be noted that, compared with traditional weak defect detection schemes that rely on mechanical scanning or speckle interference, this method achieves high-sensitivity detection of weak scratches on the surface of highly reflective workpieces by organically combining polarization imaging and multi-angle illumination, without the need for complex vibration isolation platforms and motion mechanisms, which significantly reduces the difficulty of system hardware deployment and maintenance costs.
[0059] In addition, it should be noted that the above steps can achieve pixel-level classification and post-processing of the fused feature map, accurately segment the weak scratch area and mark the geometric parameters, and visualize the detection results intuitively, thus completing the entire process from the original image to defect recognition.
[0060] In another aspect, in some embodiments, this application provides a visual inspection system for identifying surface scratches on a workpiece, with reference to... Figure 4 The figure is a schematic diagram of a visual inspection workpiece surface scratch recognition system according to some embodiments of this application. The visual inspection workpiece surface scratch recognition system includes: a data acquisition module 401, a processing module 402, and an execution module 403, which are described below: The acquisition module 401 in this application is mainly used to control the multi-angle illumination source array to sequentially illuminate the surface of the workpiece under different incident angles, and synchronously control the polarization camera to acquire multiple original images with different polarization directions at each illumination angle. Processing module 402 in this application is mainly used to determine the Stokes vector under different illumination angles based on all the original images, and extract the total intensity image, polarization degree image and polarization angle image from each Stokes vector; The processing module 402 described in this application is also used to normalize the total intensity image, polarization degree image and polarization angle image corresponding to the same illumination angle, respectively, and then perform channel superposition after making the value range consistent to obtain a multimodal feature image under the illumination angle. The processing module 402 described in this application is also used to construct a deep feature fusion network. Taking the multimodal feature images under each illumination angle as input, a convolutional neural network is used to extract the spatial feature map of each illumination angle. The response of the scratch edge in each spatial feature map is enhanced by a spatial attention mechanism. Then, all enhanced spatial feature maps are adaptively weighted and fused to generate a fused feature map. The execution module 403 in this application is mainly used to perform semantic segmentation on the fused feature map, identify and mark the weak scratch areas on the surface of the workpiece.
[0061] Each module in the aforementioned visual inspection system for identifying surface scratches on workpieces can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each module.
[0062] In another embodiment, this application provides a computer device, which may be a server, and its internal structure diagram may be as follows. Figure 5 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores data for visually inspecting and identifying surface scratches on workpieces. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a method for visually inspecting and identifying surface scratches on workpieces.
[0063] Those skilled in the art will understand that Figure 5The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0064] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described embodiment of the visual inspection workpiece surface scratch recognition method.
[0065] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps described in the above embodiment of the visual inspection method for identifying scratches on the surface of a workpiece.
[0066] In one embodiment, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps described in the embodiment of the visual inspection method for identifying scratches on a workpiece surface.
[0067] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0068] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0069] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for visually inspecting and identifying scratches on the surface of a workpiece, characterized in that, The method comprises the following steps: controlling a multi-angle illumination light source array to sequentially irradiate the surface of the detection workpiece at different incident angles, and synchronously controlling a polarization camera to collect multiple original images with different polarization directions at each illumination angle; determining Stokes vectors at different illumination angles according to all the original images, and extracting total intensity images, degree of polarization images and polarization angle images from the Stokes vectors; normalizing the corresponding total intensity images, degree of polarization images and polarization angle images at the same illumination angle to make the value ranges consistent, and then performing channel superposition to obtain a multi-modal feature image at the illumination angle; constructing a deep feature fusion network, taking the multi-modal feature images at each illumination angle as input, using a convolutional neural network to extract spatial feature maps of each illumination angle, enhancing the responses of scratch edges in the spatial feature maps through a spatial attention mechanism, and then adaptively weighting and fusing all the enhanced spatial feature maps to generate a fusion feature map; performing semantic segmentation on the fusion feature map to identify and mark the weak scratch region on the surface of the detection workpiece.
2. The method of claim 1, wherein, The determination of the Stokes vectors at different illumination angles according to all the original images specifically comprises: calculating the Stokes vector component images at each illumination angle according to the multiple original images with different polarization directions at each illumination angle; outputting the calculated Stokes vector component images as the Stokes vectors at the corresponding illumination angle.
3. The method of claim 1, wherein, The extraction of the total intensity images, the degree of polarization images and the polarization angle images from the Stokes vectors specifically comprises: extracting the total intensity image at the corresponding illumination angle from the Stokes vector at each illumination angle; calculating the degree of polarization image and the polarization angle image at the corresponding illumination angle according to the Stokes vector at each illumination angle.
4. The method of claim 1, wherein, The normalization of the corresponding total intensity images, the degree of polarization images and the polarization angle images at the same illumination angle to make the value ranges consistent, and then the channel superposition to obtain the multi-modal feature image at the illumination angle specifically comprises: normalizing the total intensity image, the degree of polarization image and the polarization angle image at the same illumination angle to make the pixel value ranges of the three consistent; performing channel superposition on the normalized total intensity image, the normalized degree of polarization image and the normalized polarization angle image to generate the multi-modal feature image at the illumination angle.
5. The method of claim 1, wherein, The extraction of the spatial feature maps of each illumination angle using a convolutional neural network specifically comprises: inputting the multi-modal feature images at each illumination angle into a convolutional neural network with shared weights; obtaining the spatial feature map corresponding to each illumination angle through forward propagation calculation of the convolutional neural network.
6. The method of claim 1, wherein, The enhancement of the responses of scratch edges in the spatial feature maps through the spatial attention mechanism specifically comprises: generating a corresponding spatial attention weight map for the spatial feature map of each illumination angle; multiplying each spatial attention weight map with the corresponding spatial feature map element by element to enhance the responses of scratch edges in the spatial feature maps.
7. The method of claim 1, wherein, The adaptive weighted fusion of all the enhanced spatial feature maps to generate a fusion feature map specifically comprises: The global description vector of the channel dimension is calculated for each enhanced spatial feature map of the illumination angle; An adaptive fusion weight of each illumination angle is learned according to each global description vector; The adaptive fusion weight is learned according to each illumination angle.
8. A visual inspection workpiece surface scratch identification system characterized by, It comprises: The acquisition module is used for controlling the multi-angle illumination light source array to irradiate the surface of the detection workpiece in turn at different incident angles, and synchronously controlling the polarization camera to collect multiple original images of different polarization directions under each illumination angle; The processing module is used for determining the Stokes vector under different illumination angles according to all original images, and extracting the total intensity image, the polarization degree image and the polarization angle image from each Stokes vector; The processing module is also used for performing normalization processing on the corresponding total intensity image, polarization degree image and polarization angle image under the same illumination angle respectively, and performing channel superposition after making the value range consistent to obtain the multi-modal feature image under the illumination angle; The processing module is also used for constructing a deep feature fusion network, taking the multi-modal feature image under each illumination angle as input, using a convolutional neural network to extract the spatial feature map of each illumination angle, enhancing the response of the scratch edge in each spatial feature map through a spatial attention mechanism, and then adaptively weighting and fusing all enhanced spatial feature maps to generate a fusion feature map; The execution module is used for performing semantic segmentation on the fusion feature map, and identifying and marking the weak scratch area on the surface of the detection workpiece. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor executes the computer program to realize the steps of the visual detection workpiece surface scratch identification method in any one of claims 1 to 7.
10. A computer readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the visual detection workpiece surface scratch identification method in any one of claims 1 to 7.