Aluminum alloy floor protection layer defect image recognition method and system

CN122335835APending Publication Date: 2026-07-03JIANGSU ZHONGTIAN ANTI-STATIC FLOORING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
JIANGSU ZHONGTIAN ANTI-STATIC FLOORING CO LTD
Filing Date
2026-05-07
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing technologies cannot effectively detect minute defects in the protective layer of aluminum alloy flooring, especially subsurface hidden defects, which leads to decreased adhesion and shortened service life of the protective layer.

Method used

The surface point cloud data is acquired through the 3D topography pre-sensing module, the local normal vector and the high reflectivity risk coefficient are calculated, the adaptive polarization control imaging module performs region division and polarization image acquisition, the multi-frame polarization fusion and defect enhancement module performs image processing, and finally the defect detection network is used to identify defects.

Benefits of technology

It enables high-precision detection of minute defects in the protective layer of aluminum alloy flooring, improving detection accuracy and reliability, and solving the problem of missed defect detection in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122335835A_ABST
    Figure CN122335835A_ABST
Patent Text Reader

Abstract

The application discloses an aluminum alloy floor protective layer defect image recognition method and system, relates to the technical field of computer vision image recognition, and comprises several function modules, including: a three-dimensional morphology pre-perception module, which obtains three-dimensional point cloud data of the surface of an aluminum alloy floor protective layer, calculates a local normal vector and a high light reflection risk coefficient, and divides a field of view into a low-curvature area and a high-curvature area to be scanned based on the high light reflection risk coefficient; an adaptive polarization regulation and control imaging module, which calculates an optimal extinction angle for the low-curvature area according to the local normal vector, fixes a detection angle, collects a single frame of polarized image, performs detection angle gradual scanning with the optimal extinction angle as the center for the high-curvature area to be scanned, and collects a plurality of polarized image sequences; and a multi-frame polarization fusion and defect enhancement module, which establishes a physical model of light intensity change with a detection angle for a plurality of polarized image sequences pixel by pixel, and separates a non-polarized light intensity component image and a polarization modulation amplitude image through fitting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision image recognition technology, specifically to a method and system for image recognition of defects in the protective layer of aluminum alloy flooring. Background Technology

[0002] Aluminum alloy flooring is widely used in aerospace, rail transportation, shipbuilding, and high-end construction due to its lightweight, high strength, and excellent corrosion resistance. To enhance its wear resistance, corrosion resistance, and decorative effect, aluminum alloy flooring is typically coated with a protective layer, such as a coating, plating, or oxide film. The quality of the protective layer directly affects the service life and operational safety of the aluminum alloy flooring, and defects in the protective layer must be rigorously inspected during the production process. Common defects in the protective layer of aluminum alloy flooring fall into two main categories: one is visible surface defects, such as scratches, bubbles, orange peel, uneven coating, and color difference; the other is subsurface hidden defects, such as delamination between the coating and the aluminum alloy substrate. These defects are difficult to detect under conventional visible light imaging but significantly reduce the adhesion of the protective layer, potentially leading to large-scale peeling during service.

[0003] In existing technologies, due to the significant specular reflection characteristics of both the aluminum alloy substrate and the protective layer, fixed-angle orthogonal polarization extinction or diffuse lighting is often used to suppress reflection. However, the surface of the aluminum alloy floor protective layer is not an ideal plane; its roller coating texture and local undulations cause differences in the surface normal vectors at various locations. Fixed polarization angle schemes cannot achieve optimal extinction for different normal micro-elements simultaneously, resulting in severe specular reflection residue in some areas, and tiny defects are completely submerged in the high-brightness spot. Existing polarization image processing methods only involve simple difference or polarization degree calculation of images with different polarization angles, without deeply utilizing the defect edge information contained in the gradual change of polarization state with surface morphology. This leads to blurred defect edges and loss of details in the image after extinction. At the same time, the single visible light imaging mode lacks the ability to perceive subsurface hidden defects. Defects such as delamination and debonding between the protective layer and the aluminum alloy substrate cannot be reflected by visible light images, resulting in systematic missed detection of such defects and an inability to fully evaluate the adhesion quality of the protective layer. Summary of the Invention

[0004] To achieve the above objectives, the present invention provides the following technical solution: An image recognition system for defects in aluminum alloy floor protective layers includes: The 3D shape pre-sensing module acquires 3D point cloud data of the surface of the aluminum alloy floor protective layer, calculates the local normal vector and high reflectivity risk coefficient, and divides the field of view into a low curvature region and a high curvature region to be scanned based on the high reflectivity risk coefficient. The adaptive polarization control imaging module calculates the optimal extinction angle for low curvature regions based on local normal vectors and acquires a single-frame polarization image with a fixed polarization angle. For high curvature regions to be scanned, it performs a gradual polarization angle scan centered on the optimal extinction angle and acquires a multi-frame polarization image sequence. The multi-frame polarization fusion and defect enhancement module establishes a physical model of light intensity variation with the detection angle for each pixel of the multi-frame polarization image sequence. By fitting, it separates the non-polarized light intensity component image and the polarization modulation amplitude image. The single-frame polarization image is used as the non-polarized light intensity component in the low curvature region and synthesized with the non-polarized light intensity component image in the high curvature region to form a full-field non-polarized base image. The module uses the local gradient ratio of the two channels to construct adaptive weights and weightedly fuses the polarization modulation amplitude image into the full-field non-polarized base image to generate a defect enhancement image. The defect recognition module inputs the enhanced defect image into a pre-trained defect detection network and outputs the defect category, location, and confidence level.

[0005] Furthermore, the process of calculating the local normal vector and the high reflectivity risk coefficient is as follows: Based on pixel-by-pixel analysis of 3D point clouds, depth gradients are derived. Normal vectors are constructed using negative horizontal and vertical gradient values ​​and normalized. After mean filtering, a smoothed normal vector field is obtained. Based on the smoothed normal vector field, the rate of change of each component of the normal vector with respect to the horizontal and vertical coordinates is obtained. The sum of the horizontal and vertical moduli of change is taken as the high reflectivity risk coefficient. The high reflectivity risk coefficient is compared with a preset threshold to divide the low curvature region and the high curvature region to be scanned.

[0006] Furthermore, the process of dividing the field of view into a low-curvature region and a high-curvature region to be scanned is as follows: The high reflectivity risk coefficient of each pixel is compared with a preset threshold to determine high curvature pixels and low curvature pixels. Connectivity analysis is performed on high curvature pixels, and isolated connected components with an area smaller than the preset threshold are removed to generate a binary region label map. High curvature regions are assigned a value of 1, and low curvature regions are assigned a value of 0.

[0007] Furthermore, the process of calculating the optimal extinction angle for the low curvature region based on the local normal vector and acquiring a single-frame polarization image with a fixed analyzer angle is as follows: Read the low curvature region marked with a value of 0 in the region marker map and extract the smooth normal vector within the region; calculate the average value of the vertical deflection angle of the normal vector within the region as the equivalent tilt angle, calculate the average value of the arctangent of the ratio of the vertical component to the horizontal component of the normal vector as the equivalent polarization rotation angle, and superimpose the base extinction angle with the equivalent polarization rotation angle to obtain the optimal extinction angle for the region; adjust the analyzer to the optimal extinction angle and fix the starting angle, and trigger the camera to acquire a single frame polarization image.

[0008] Furthermore, the process of performing a gradual scan of the high-curvature scanning area with the optimal extinction angle as the center is as follows: The median of the optimal extinction angle distribution of each pixel in the high curvature connected region is taken as the reference optimal extinction angle; the scanning range is extended to both sides by a preset multiple standard deviation with the reference optimal extinction angle as the center, and the scanning step size is determined according to the polarization extinction half-width angle; an equal arithmetic polarization detection angle scanning sequence is generated with the reference optimal extinction angle as the center; the starting angle is fixed, the polarization detection angle is adjusted frame by frame according to the scanning sequence and polarization images are acquired, and the multi-frame image sequence and the corresponding polarization detection angle sequence are packaged and output.

[0009] Furthermore, the process of separating the unpolarized light intensity component image and the polarization modulation amplitude image through fitting is as follows: N frames of light intensity vectors and corresponding polarization detection vectors are extracted pixel by pixel; the light intensity model is linearly expanded into a combination of three basis functions, and an N×3 design matrix is ​​constructed. The three combination coefficients are solved at once by multiplying the pre-stored inverse matrix with the light intensity vector; the unpolarized light intensity component is obtained by subtracting the polarization modulation amplitude from the first combination coefficient, and the square root of the sum of the squares of the second and third combination coefficients is taken as the polarization modulation amplitude; the fitting residual between the measured light intensity value and the model fitting value is calculated pixel by pixel. When the fitting residual exceeds the preset residual tolerance threshold, the polarization response of the pixel is determined to deviate from the ideal model. The minimum value of the light intensity sequence is taken as the unpolarized light intensity of the pixel, and the polarization modulation amplitude is set to the preset minimum value.

[0010] Furthermore, the process of generating defect-enhanced images is as follows: Calculate the local gradient magnitude of the unpolarized light intensity component image and the polarization modulation amplitude image, calculate the ratio of their gradient magnitudes pixel by pixel and normalize it to construct an adaptive weight map; multiply the polarization modulation amplitude image pixel by pixel by the adaptive weight map and the preset global enhancement coefficient, and then superimpose it pixel by pixel with the unpolarized light intensity component image. After dynamic range cropping and grayscale stretching, a defect enhancement image is generated.

[0011] Furthermore, the process of outputting the defect category, location, and confidence level is as follows: A defect detection network is adopted, consisting of a backbone network, a feature pyramid, and a detection head cascaded together. The enhanced defect image is input into the network, and after multi-scale feature extraction and fusion, the region proposal network generates candidate regions. The classification and regression sub-network outputs the defect category probability and bounding box parameters. After non-maximum suppression and confidence threshold filtering, the defect category, location coordinates, and confidence level are output.

[0012] A method for image recognition of defects in the protective layer of aluminum alloy flooring includes the following steps: Step 1: Obtain the 3D point cloud data of the aluminum alloy floor protective layer surface, calculate the local normal vector and the high reflectivity risk coefficient, and divide the field of view into a low curvature region and a high curvature region to be scanned based on the high reflectivity risk coefficient. Step 2: Calculate the optimal extinction angle for the low curvature region based on the local normal vector and fix the polarization angle to acquire a single-frame polarization image. Perform a gradual polarization angle scan on the high curvature region to be scanned with the optimal extinction angle as the center, and acquire a multi-frame polarization image sequence. Step 3: Establish a physical model of light intensity variation with polarization angle for each pixel of the multi-frame polarization image sequence. Separate the non-polarized light intensity component image and the polarization modulation amplitude image by fitting. Use the single-frame polarization image as the non-polarized light intensity component in the low curvature region and synthesize it with the non-polarized light intensity component image in the high curvature region to form a full-field non-polarized base image. Construct adaptive weights using the dual-channel local gradient ratio and weightedly fuse the polarization modulation amplitude image into the full-field non-polarized base image to generate a defect enhancement image. Step 4: Input the enhanced image of the defect into the pre-trained defect detection network, and output the defect category, location, and confidence level.

[0013] The present invention provides a method and system for image recognition of defects in the protective layer of aluminum alloy flooring, which has the following beneficial effects: (1) This invention obtains surface point clouds and calculates local normal vectors and high reflectivity risk coefficients through a three-dimensional topography pre-sensing module, and drives an adaptive polarization control imaging module to fix the optimal extinction angle in low curvature areas and perform a gradual polarization angle scan in high curvature areas: so that the polarization imaging parameters are dynamically matched with the surface topography, instead of using a fixed polarization angle, fundamentally solving the problem that the existing orthogonal polarization extinction method cannot adapt to the non-uniform surface of the aluminum alloy floor protective layer and the defects in high reflectivity areas are submerged by specular reflection.

[0014] (2) This invention establishes a physical model of light intensity changing with the detection angle pixel by pixel through multi-frame polarization fusion and defect enhancement module. The non-polarized light intensity component image and polarization modulation amplitude image are separated by least squares fitting. The adaptive weight is constructed using the local gradient ratio of the two channels for weighted fusion. This effectively removes the specular reflection component while selectively enhancing the polarization details of the defect edge. This solves the problem that existing polarization separation methods only perform simple difference or polarization degree calculation and cannot use polarization gradient information to enhance small defects.

[0015] (3) The present invention controls the polarization imaging parameters by using the normal vector field and region marker map output by the three-dimensional topography pre-sensing module, and drives the fitting separation and fusion enhancement by using the multi-frame polarization sequence output by the adaptive polarization control imaging module. The detection results of the defect recognition module are fed back to optimize the imaging and fusion parameters. This closed-loop collaborative mechanism makes the surface topography perception, polarization optical control and intelligent defect recognition interdependent and mutually optimized, and realizes the overall improvement of the detection accuracy of small defects on highly reflective surfaces. Attached Figure Description

[0016] Figure 1 This is a schematic diagram of the system of the present invention; Figure 2 This is a schematic diagram of the overall method of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example 1: Please see Figure 1 Embodiment 1 of this application provides an image recognition system for defects in the protective layer of aluminum alloy flooring. The system includes: The three-dimensional shape pre-sensing module acquires three-dimensional point cloud data of the surface of the aluminum alloy floor protective layer, calculates the local normal vector and high reflectivity risk coefficient, and divides the field of view into a low curvature region and a high curvature region to be scanned. Obtain 3D point cloud data of the surface of the aluminum alloy floor protective layer: A structured light 3D sensor projects a multi-frequency phase-shift coded stripe pattern onto the surface of an aluminum alloy floor protective layer, simultaneously acquiring deformed stripe images modulated by the surface morphology. Phase calculation is performed on the deformed stripe images, and the absolute phase value of each pixel is extracted using a phase-shift algorithm. Based on pre-calibrated triangulation parameters, the absolute phase values ​​are converted into spatial depth coordinates to generate dense 3D point cloud data characterizing the surface morphology. Median filtering is applied to the point cloud data to remove outlier noise points, resulting in a smooth 3D point cloud.

[0019] Specifically, a structured light 3D sensor is installed directly above the inspection station. The sensor consists of a digital light processing projection unit and a high-speed industrial camera, with their optical axes fixed at a preset angle on a precision support. The projection unit projects a set of multi-frequency phase-shift coded stripe patterns onto the surface of the aluminum alloy floor protective layer. The frequencies are selected from four bands: level 1, level 4, level 16, and level 64. At each frequency, four sinusoidal stripe images with phase shifted by 90° are projected, for a total of 16 patterns. Simultaneously, the industrial camera is triggered to acquire a sequence of deformed stripe images modulated by the surface morphology of the protective layer. The exposure time is set to 0.5 milliseconds to adapt to the high reflectivity of the aluminum alloy surface and avoid overexposure.

[0020] Pixel-by-pixel phase calculation is performed on four phase-shifted images at each frequency. Let the relative phase shifts of the four images be 0°, 90°, 180° and 270°, respectively, and the corresponding pixel gray values ​​be denoted as gray level 1, gray level 2, gray level 3 and gray level 4, respectively. Then the wrapping phase value at this frequency is calculated by a four-step phase shift algorithm. Its value is the arctangent of the difference between gray level 4 and gray level 2 divided by the difference between gray level 1 and gray level 3. The obtained wrapping phase is restricted to the main value range of -π / 2 to +π / 2 and has periodic jumps.

[0021] To eliminate phase ambiguity, a multi-frequency heterodyne method is used to expand the phase of the wrap-around phase at four frequencies. First, the phase of the first-level frequency is taken as the reference, and its phase is unique and has no jumps throughout the field. Then, the frequency ratio between the fourth-level frequency and the first-level frequency is used to expand the phase level by level, and the sub-integer of the stripe level of the corresponding pixel is solved. Finally, the full-field continuous absolute phase distribution map at 64 frequencies is obtained.

[0022] Based on pre-calibrated system parameters, including the baseline distance between the projection unit and the camera, the optical axis angle, and the camera intrinsic parameter matrix, the absolute phase value is converted into a depth value through triangulation. Assuming that the vertical stripes projected by the projection unit produce a linear phase change along the horizontal direction, there is a linear mapping relationship between the phase difference and the depth offset. By looking up the pre-calibrated phase-depth mapping table, the depth offset of the point relative to the reference plane is obtained pixel by pixel. After superimposing the height of the reference plane, the spatial depth coordinates of the object point are obtained.

[0023] By combining the camera pixel coordinate system with the depth value, we obtain 3D point cloud data with the camera optical center as the origin. Each valid pixel corresponds to a spatial point in the point cloud, and the data format is a triplet of horizontal coordinate, vertical coordinate and depth coordinate.

[0024] The initial 3D point cloud is subjected to bilateral filtering. The filter kernel considers both spatial proximity and depth value similarity. The standard deviation in the spatial domain is set to 1 pixel, and the standard deviation in the depth domain is set to 0.5 mm. The bilateral filter effectively suppresses depth jitter caused by sensor thermal noise and shot noise while maintaining the defect edge characteristics. After filtering, outliers with depth values ​​exceeding the effective range are removed, and sparse missing areas are repaired by linear interpolation of the depth values ​​of adjacent effective points. Finally, smooth and complete dense 3D point cloud data is generated, covering the entire detection area of ​​the aluminum alloy floor protective layer.

[0025] The process of calculating the local normal vector and the high reflectivity risk coefficient is as follows: For each valid pixel in the 3D point cloud, the first partial derivatives of its depth value with respect to the horizontal and vertical coordinates are taken. The horizontal gradient is obtained by subtracting the left depth value from the right depth value of the pixel and dividing by twice the pixel spacing. The vertical gradient is obtained by subtracting the upper depth value from the lower depth value and dividing by twice the pixel spacing. For pixels at the point cloud boundary where it is not possible to take bilateral differences, a one-sided forward or backward difference approximation is used. This yields two gradient fields, which describe the rate of change of the surface height in the horizontal and vertical directions, respectively.

[0026] For each pixel location, the negative value of the horizontal gradient is used as the first component of the normal vector, the negative value of the vertical gradient is used as the second component of the normal vector, and the constant 1 is used as the third component of the normal vector, forming an unnormalized initial normal vector. This construction method transforms the gradient information into a normal representation of the surface tangent plane. The Euclidean norm of this initial normal vector is calculated, which is the square root of the sum of the squares of the three components. Each component is divided by this norm to obtain the unit normal vector. The direction of the unit normal vector represents the spatial orientation of the local surface element at the pixel location.

[0027] Mean filtering is performed on each component of the unit normal vector field. A 3x3 neighborhood is selected as the filtering window, and the weights of each pixel in the window are equal. The arithmetic mean of the normal vector components in the neighborhood is used to replace the corresponding component of the center pixel. After filtering, the normal vector of each pixel is normalized again to eliminate the length deviation caused by the mean operation and obtain a smooth normal vector field. The smoothing process suppresses the interference of high-frequency noise from the sensor on the normal vector estimation, while preserving the macroscopic undulation characteristics of the surface.

[0028] For each pixel in the smoothed normal vector field, calculate the rate of change of its horizontal component and the rate of change of its vertical component with respect to the horizontal coordinate. Take the square root of the sum of their squares to obtain the horizontal normal vector change modulus. Similarly, calculate the rate of change of its horizontal component and the rate of change of its vertical component with respect to the vertical coordinate, and take the square root of their squares to obtain the vertical normal vector change modulus. Add the horizontal and vertical normal vector change moduli together, and the sum is the high reflectivity risk coefficient for that pixel location. This coefficient is essentially the sum of the moduli of the normal vector space gradient, reflecting the magnitude of the local surface curvature. The greater the curvature, the more severe the surface bending, the greater the gradient of the specular reflection direction, and the higher the high reflectivity risk.

[0029] The process of dividing the low-curvature region and the high-curvature region to be scanned is as follows: A high curvature risk threshold is preset. This threshold is determined by calibration based on the correspondence between the typical texture undulation amplitude of the aluminum alloy floor protective layer surface and the polarization scanning angle range. The calibration method is as follows: select calibration samples with different curvature distributions, evaluate the specular reflection residual amount in each area after imaging with a fixed polarization angle, and use the minimum risk coefficient corresponding to the area where the specular reflection residual amount exceeds the preset tolerance as the classification threshold. The typical threshold value ranges from 0.02 mm to 0.1 mm.

[0030] The system iterates through all valid pixels within the field of view covered by the 3D point cloud; it compares the calculated high reflectivity risk coefficient of each pixel location with the high curvature risk threshold one by one; if the high reflectivity risk coefficient of a pixel is greater than or equal to the threshold, the pixel is determined to belong to the high curvature region; if the risk coefficient is less than the threshold, the pixel is determined to belong to the low curvature region.

[0031] For pixels belonging to high curvature regions in the judgment results, connected component analysis is performed. The 8-neighbor connectivity criterion is used to merge spatially adjacent high curvature pixels into independent regions. The pixel area of ​​each connected component is counted, and isolated connected components with an area smaller than the preset minimum area threshold are reclassified as low curvature regions to eliminate scattered misjudged points caused by local noise in the 3D point cloud. The minimum area threshold is set according to the camera spatial resolution and the minimum identifiable defect size, with a typical value of 5×5 pixel area.

[0032] The connected components of the high-curvature region and the remaining low-curvature region are numbered and labeled to generate a binary region label map of the same size as the original point cloud. The high-curvature region to be scanned in the label map is assigned a value of 1, and the low-curvature region is assigned a value of 0. The label map indicates the polarization imaging strategy to be adopted for the corresponding spatial position pixel by pixel. The high-curvature region performs a gradual polarization angle scan, and the low-curvature region performs a fixed polarization angle single-frame acquisition. The label map serves as the control input of the adaptive polarization control imaging module to realize the field-of-view partition differential imaging decision.

[0033] The adaptive polarization control imaging module calculates the optimal extinction angle for low curvature regions based on local normal vectors and acquires a single-frame polarization image with a fixed polarization angle. For high curvature regions to be scanned, it performs a gradual polarization angle scan centered on the optimal extinction angle and acquires a multi-frame polarization image sequence. The process of acquiring a single-frame polarization image is as follows: The system receives region marker maps from the 3D shape pre-sensing module, extracts the pixel positions of all low-curvature regions with a value of 0, and determines the spatial distribution range of low-curvature regions in the field of view. It then processes each low-curvature connected region separately and extracts the smoothed normal vector data corresponding to each pixel in that region.

[0034] For the currently processed low-curvature connected region, traverse the local normal vectors of all pixels in the region and calculate the vertical deflection angle of each normal vector; the vertical deflection angle is defined as the angle between the normal vector and the vertical direction, reflecting the degree of inclination of the surface of the micro element relative to the horizontal plane; since the region is a low-curvature area, the direction of the normal vector in the region changes gently, and the arithmetic mean of the vertical deflection angles of all normal vectors in the region is used as the equivalent tilt angle of the region.

[0035] According to the optical path geometry of coaxial polarization imaging, when the surface element is tilted, the relative angle between the incident light and the surface normal shifts, causing the polarization direction of the specular reflection to rotate. The rotation angle is related to the projection direction of the normal vector in the horizontal plane and is approximately equal to the arctangent of the ratio of the vertical component to the horizontal component of the normal vector. The arithmetic mean of this rotation angle of all pixels in the domain is taken as the equivalent polarization rotation angle of the region.

[0036] The optimal extinction angle of the low-curvature connected domain is obtained by superimposing the system's preset basic extinction angle with the region's equivalent polarization rotation angle; the basic extinction angle is the light source's polarization angle plus 90°.

[0037] The embedded processing unit sends control commands to the electrically controlled rotating analyzer of the polarization camera, rotating the analyzer along its transmission axis to the calculated optimal extinction angle position. The rotation is driven by a closed-loop stepper motor, with an angle positioning accuracy better than 0.1°. After the analyzer is in place, a preset stabilization time is delayed to allow mechanical vibration to decay, ensuring angle stability.

[0038] The electronically controlled polarizer of the synchronously controlled polarized light source keeps the polarization angle fixed at the system preset value; the light source adopts a pulse trigger mode and is synchronously lit within the camera exposure window, with the width of a single pulse being consistent with the camera exposure time, providing sufficient illumination intensity while avoiding the accumulation of thermal effects on the surface of the protective layer.

[0039] The polarization camera is triggered to perform single-frame exposure acquisition of the low-curvature connected domain. The camera exposure time is adaptively adjusted according to the reflectivity of the protective layer surface to stabilize the average gray value of the image within the preset target range. In the acquired single-frame polarization image, the specular reflection component is suppressed to the maximum extent under the optimal extinction angle, and the diffuse reflection component becomes the main component of the image.

[0040] When multiple separate low-curvature connected domains exist within the field of view, the optimal extinction angle of each domain is calculated one by one and single-frame polarization images are acquired sequentially; the images of each domain are stitched together according to their spatial coordinates to restore a complete low-curvature region polarization image; overlapping pixels at the boundaries of adjacent connected domains are directly filled with the gray values ​​of the corresponding low-curvature domain images.

[0041] The acquired single-frame polarization image of the complete low curvature region is output to the multi-frame polarization fusion and defect enhancement module. This image serves as an approximation of the non-polarized light intensity component in the low curvature region. Together with the non-polarized light intensity component image obtained by fitting and separating the high curvature region, it constitutes the diffuse reflection substrate data of the entire field of view, which is then processed uniformly in the subsequent fusion and enhancement steps.

[0042] The process of performing a gradual scan with a detection angle is as follows: The region marker map is received from the 3D topography pre-sensing module, and the pixel positions of all high-curvature scanned regions with a value of 1 are extracted. Each independent high-curvature connected region is processed separately, and the smoothed normal vector data corresponding to each pixel in the region is loaded. The surface normal vector space of the high-curvature region changes drastically, and a single extinction angle cannot simultaneously meet the optimal extinction requirements of each micro-element in the region. Therefore, a complete polarization response sequence needs to be obtained by polarization angle scanning.

[0043] Traverse the local normal vectors of all pixels in the current high-curvature connected region, and calculate the optimal extinction angle for each micro-element pixel by pixel using the same method as for low-curvature regions; statistically analyze the distribution histogram of the optimal extinction angles of all pixels in the region, and take the median of the distribution as the reference optimal extinction angle for the region; the median is robust to outliers and can effectively represent the extinction demand center of most micro-elements in the region, and is used as the scanning center angle.

[0044] The scanning angle range is determined based on the degree of normal vector dispersion in the high curvature region; the standard deviation of the optimal extinction angle of each pixel in the domain is calculated, and the standard deviation of the reference optimal extinction angle is extended to both sides by a preset multiple as the upper and lower limits of scanning; the preset multiple is calibrated based on the typical curvature of the protective layer surface and the polarization extinction angle sensitivity curve, usually taking 2 to 3 times the standard deviation to ensure that at least 95% of the pixels in the coverage area have the optimal extinction angle distribution range.

[0045] The scanning step size is determined based on the analyzer's angular resolution and the polarization extinction half-width angle. The polarization extinction half-width angle is defined as the angular offset when the light intensity increases from the optimal extinction angle to twice the minimum extinction value, and is obtained through calibration measurement. The scanning step size cancels half of the light intensity half-width angle to ensure that distinguishable light intensity changes can be captured at any pixel position between two adjacent frames. The typical step size ranges from 2° to 5°.

[0046] The total number of scan frames is determined by dividing the scan range by the step size, rounding up and adding 1. The frame number starts from 1 and is sequentially incremented.

[0047] Using the optimal extinction angle as the scanning center, calculate the detection angle setting value corresponding to each frame in the scanning sequence; the detection angle of the first frame is equal to the scanning center angle minus half of the scanning range, and the detection angle of each subsequent frame is increased by one step based on the previous frame, and the detection angle of the last frame is equal to the scanning center angle plus half of the scanning range; all detection angle setting values ​​form an arithmetic sequence, covering the entire scanning range.

[0048] The electrically controlled polarizer of the polarizing light source is fixed at the preset polarization angle position of the system and remains unchanged throughout the scanning process; the light source adopts a pulse triggering mode synchronized with the camera frame rate. For each frame of image acquired by the camera, the light source is synchronously lit once within the exposure window, and the pulse width is equal to the camera exposure time; the output light intensity of the light source remains constant during the scanning process, and the light intensity is stabilized through closed-loop feedback of the light source controller.

[0049] The embedded processing unit sends angle control commands to the electronically controlled rotating analyzer in sequence according to the analyzer angle settings corresponding to each frame in the scanning sequence. After the analyzer rotates to the correct position, a preset stabilization time is delayed to eliminate residual mechanical vibration. Then, a trigger signal is sent to make the polarization camera acquire one frame of image and record the actual value of the analyzer angle corresponding to that frame. After acquisition, the analyzer rotates to the next set angle and repeats the above process until all set angles in the scanning sequence have been traversed. During the acquisition of each frame of image, parameters such as camera exposure time and gain remain constant to ensure that only the analyzer angle changes between frames.

[0050] Acquire a sequence of multiple polarization images: The complete multi-frame polarization image sequence acquired from the current high curvature connected domain is arranged by frame number, with the corresponding polarization detection angle sequence value for each frame, and output to the multi-frame polarization fusion and defect enhancement module in the form of a data set. The data set contains metadata such as image sequence, angle sequence, scanning center angle, scanning range, step size, and frame number, which are used for subsequent pixel-by-pixel polarization response fitting. When there are multiple high curvature connected domains in the field of view, each domain independently performs the above scanning process and outputs the corresponding multi-frame polarization image sequence data set.

[0051] The multi-frame polarization fusion and defect enhancement module establishes a physical model of light intensity variation with polarization angle for each pixel of the multi-frame polarization image sequence. By fitting, it separates the non-polarized light intensity component image and the polarization modulation amplitude image, and uses the local gradient ratio of the two channels to construct adaptive weights to weightedly fuse the polarization modulation amplitude image into the non-polarized light intensity component image to generate a defect enhancement image. Establish a physical model of how light intensity varies with the analyzer angle: For a multi-frame polarization image sequence acquired in a high curvature region, the light intensity value of each pixel in each frame is extracted using the pixel spatial position as the index. The light intensity values ​​are then arranged into a light intensity sequence according to the frame number, and the sequence length is equal to the total number of scanned frames. At the same time, the measured values ​​of the polarization angles corresponding to the light intensity sequence of each pixel are extracted and arranged into an angle sequence in the same order. The two sequences correspond one-to-one, forming a set of data pairs to be fitted.

[0052] Based on the principles of polarization optics, under coaxial illumination and a fixed polarization angle, the polarization behavior of the light reflected from the protective layer surface can be analytically described. The specular reflection component maintains the linear polarization state of the incident light, and its intensity after passing through the analyzer follows Malus's law, being proportional to the square of the cosine of the angle between the analyzer and the specular reflection polarization direction. The diffuse reflection component depolarizes due to surface scattering and appears as unpolarized light. Its intensity after passing through the analyzer is independent of the analyzer angle and remains constant. Therefore, the total pixel intensity can be expressed as the sum of two parts: an unpolarized substrate intensity independent of the analyzer angle, plus a polarized light intensity component modulated by the analyzer angle.

[0053] By transforming trigonometric identities, the cosine square form of Malus's law is rewritten into a standard cosine function form, yielding a physical model structure for the variation of pixel light intensity with the detection angle: the model consists of three terms. The first term is a constant term, representing the average light intensity level; the second term is a cosine oscillation term, whose amplitude characterizes the polarization modulation depth, and whose phase reflects twice the equivalent optimal extinction angle of the pixel; the third term is implicit in the phase and amplitude, and the overall performance is that the light intensity fluctuates cosinely with the detection angle, and the fluctuation frequency is twice the frequency of the detection angle change.

[0054] The model contains three undetermined parameters: constant parameter, modulation amplitude parameter, and equivalent phase parameter. The difference between the constant parameter and the modulation amplitude parameter is the unpolarized light intensity component of the pixel, the modulation amplitude parameter is the polarization modulation amplitude of the pixel, and half of the equivalent phase parameter is the equivalent optimal extinction angle of the pixel.

[0055] For each pixel, the light intensity sequence and angle sequence are obtained, and least squares curve fitting is performed using the aforementioned physical model. To transform the fitting into a linear problem, the model expression is expanded into a three-term linear combination: the first term is 1 multiplied by a constant term, the second term is twice the cosine value of the detection angle multiplied by a combination parameter, and the third term is twice the sine value of the detection angle multiplied by another combination parameter. This yields a set of linear observation equations for the three combination parameters. The number of equations is equal to the number of scan frames, and there are three parameters to be determined.

[0056] An overdetermined system of equations is constructed, and the solution to its normal equations is obtained by the least squares method. The coefficient matrix of the normal equations is only related to the detection angle of each frame. Under a constant scanning angle sequence and frame number, it is a fixed matrix, and its inverse matrix can be pre-calculated and stored. For each pixel, the light intensity observation vector is multiplied by the pre-stored inverse matrix, and the estimated values ​​of the three combined parameters can be obtained in one operation.

[0057] Physical parameters are further extracted from the three combined parameters: the constant parameter is the first combined parameter; the polarization modulation amplitude parameter is the square root of the sum of the squares of the second and third combined parameters; and the equivalent optimal extinction angle is half the arctangent of the ratio of the second and third combined parameters.

[0058] Based on the constraint relationship between the physical model parameters, the unpolarized light intensity component value is calculated pixel by pixel. Its value is the constant term parameter minus the polarization modulation amplitude parameter. This value represents the equivalent light intensity of diffuse reflection after removing the specular reflection polarization component, which constitutes the unpolarized light intensity component image. The polarization modulation amplitude parameter directly constitutes the polarization modulation amplitude image, reflecting the sensitivity of each pixel to polarization state modulation. The two images are spatially registered and have the same size, serving as the data basis for subsequent gradient weighted fusion.

[0059] Residual analysis is performed pixel by pixel on the fitting results to calculate the root mean square error between the fitted value and the measured value of the light intensity of the pixel. If the root mean square error exceeds the preset noise tolerance, it is determined that the polarization response of the pixel deviates from the ideal model, which may be caused by local defects. At this time, the pixel is marked as a fitting outlier, and its unpolarized light intensity is replaced by the minimum light intensity value in the original sequence. The polarization modulation amplitude is set to the preset minimum value to ensure the integrity of the image. Finally, the verified unpolarized light intensity component image and polarization modulation amplitude image are output.

[0060] The process of separating the unpolarized light intensity component image from the polarization modulation amplitude image is as follows: For a multi-frame polarization image sequence acquired in a high curvature region, the number of image frames is N, and the spatial resolution of each image frame is W columns × H rows. Using pixel coordinates as indices, N light intensity values ​​at each pixel location in the N images are extracted and arranged into a light intensity vector according to the frame number. Simultaneously, N measured polarization angle values ​​corresponding to the pixel location are extracted and arranged into an angle vector according to the same frame number. The length of both the light intensity vector and the angle vector is N. Data at the same index position of the two vectors constitute an observation data pair, for a total of N pairs, which are used as the fitting data input for that pixel.

[0061] Based on the cosine response law of light intensity varying with the detection angle in polarization optics, a physical model of pixel light intensity with respect to the detection angle is constructed. The model expresses light intensity as the sum of a constant term and a cosine oscillation term, with the cosine oscillation term taking twice the detection angle as the variable. Through trigonometric identity expansion, the model is transformed into a linear combination of three basis functions: the first basis function is a constant 1, the second basis function is the cosine value of twice the detection angle, and the third basis function is the sine value of twice the detection angle. The coefficients corresponding to the three basis functions are as follows: the first coefficient is the average light intensity parameter, the second coefficient is the product of the modulation amplitude and twice the equivalent phase cosine value, and the third coefficient is the product of the modulation amplitude and twice the equivalent phase sine value.

[0062] For each pixel's N angle values, calculate the values ​​of the corresponding three basis functions to construct a design matrix. The design matrix has N rows and 3 columns. The first column contains all 1s, the second column contains the cosine of twice the angle of each detection angle, and the third column contains the sine of twice the angle of each detection angle.

[0063] Using the constructed design matrix and the extracted light intensity vector, a linear observation equation system is established, with the number of equations N being greater than the number of unknowns (3). The least squares method is used to solve for the optimal estimates of the coefficients of the three basis functions.

[0064] To solve the problem, first calculate the product of the transpose of the design matrix and itself to obtain a 3rd order square matrix; then calculate the product of the transpose of the design matrix and the light intensity vector to obtain a 3-dimensional column vector; solve the 3rd order linear equation system to obtain three coefficient estimates, which are denoted as the first combination coefficient, the second combination coefficient and the third combination coefficient, respectively.

[0065] Since the design matrix depends only on the deflection angle sequence, and the scanning angle sequence for the same high curvature region is the same for all pixels, the design matrix and the inverse matrix of its derived 3rd order square matrix are uniform constants that can be pre-calculated and stored at once. The solution operation for each pixel is simplified to multiplying the pre-stored inverse matrix with the light intensity vector of that pixel to directly obtain the estimated values ​​of the three combination coefficients. The pixel-by-pixel operation process is iterative and suitable for large-scale parallel execution on a graphics processor.

[0066] The physical parameters are derived analytically from three combination coefficients, and the specific calculation rules are as follows: The unpolarized light intensity component is equal to the first combination coefficient minus the polarization modulation amplitude; the polarization modulation amplitude is equal to the square root of the sum of the squares of the second and third combination coefficients; the above extraction operation is performed pixel by pixel, and the unpolarized light intensity component values ​​are arranged according to the pixel spatial position to generate an unpolarized light intensity component image; the size of the unpolarized light intensity component image is W columns × H rows, and each pixel value is a floating-point gray value, which physically means that only the diffuse reflection equivalent light intensity remains after removing the specular reflection polarization component.

[0067] The polarization modulation amplitude values ​​are arranged according to the spatial position of pixels to generate a polarization modulation amplitude image. The polarization modulation amplitude image has the same size of W columns × H rows. Each pixel value reflects the intensity of the polarization component in the reflected light at that position. Higher amplitude values ​​are usually observed at defect edges and places where the coating thickness changes abruptly.

[0068] The fitting residual is calculated pixel by pixel. The residual is defined as the square root of the sum of the squares of the differences between the N measured light intensity values ​​and the corresponding fitted light intensity values ​​of the pixel. The fitted light intensity value is obtained by substituting the three combined coefficients obtained from the solution into the model expression frame by frame.

[0069] A tolerance threshold for the fitting residual is set, which is determined comprehensively based on the standard deviation of camera noise and the polarization fluctuation amplitude caused by the normal texture of the protective layer surface. The fitting residual of each pixel is compared with the tolerance threshold. Pixels with residuals less than the threshold are marked as normal fitting points, and the previously extracted non-polarized light intensity component and polarization modulation amplitude are directly used as output values. Pixels with residuals greater than or equal to the threshold are marked as fitting outliers, indicating that the polarization response at this position deviates from the ideal cosine model, which may be caused by depolarization anomalies caused by local strong defects. Such outliers are handled by a replacement strategy: the non-polarized light intensity is replaced by the minimum value in the light intensity sequence of N frames for that pixel, and the polarization modulation amplitude is replaced by a preset minimum constant value.

[0070] The unpolarized light intensity component image and the polarization modulation amplitude image, after verification processing, are output in a paired manner to the subsequent fusion enhancement stage. The two images are spatially fully registered, with each pixel corresponding to the next, and together they constitute a full-field polarization separation result dataset, providing the input data basis for subsequent adaptive weighted fusion based on the dual-channel gradient ratio.

[0071] The process of constructing adaptive weights using the two-channel local gradient ratio is as follows: The local gradient is calculated pixel-by-pixel for the unpolarized light intensity component image. The horizontal gradient is obtained by subtracting the gray value of the left adjacent pixel from the gray value of the right adjacent pixel and then dividing by twice the pixel spacing. The vertical gradient is obtained by subtracting the gray value of the upper adjacent pixel from the gray value of the lower adjacent pixel and then dividing by twice the pixel spacing. A one-sided difference approximation is used for the image boundary pixels. The gradient vector is synthesized from the horizontal and vertical gradient components, and the Euclidean norm of the vector is calculated, which is the square root of the sum of the squares of the horizontal and vertical components. This is taken as the unpolarized light intensity gradient magnitude of the pixel. The gradient magnitude reflects the degree of spatial variation of diffuse light intensity at that location. The gradient magnitude increases significantly at the edge of defects and contours, and approaches 0 in uniform texture areas.

[0072] The same gradient calculation method is used for the polarization modulation amplitude image. The horizontal gradient component and the vertical gradient component are obtained pixel by pixel. The gradient vector is synthesized and its Euclidean norm is calculated to obtain the polarization modulation amplitude gradient magnitude of the pixel. The polarization modulation amplitude image records the intensity of the polarization component in the reflected light of each pixel. Its gradient magnitude is relatively high in regions where the polarization characteristics change abruptly, such as coating thickness steps, material interfaces, and stress concentration edges, but it is usually lower than the intensity gradient magnitude of unpolarized light at normal surface textures.

[0073] The ratio of the intensity gradient magnitude of unpolarized light to the amplitude gradient magnitude of polarization modulation is calculated pixel by pixel. To avoid an infinitely large ratio due to a denominator of 0, a very small positive constant is added to the denominator. This constant is taken as 1‰ of the average global gradient magnitude of the polarization modulation amplitude image. The ratio results constitute a dual-channel local gradient ratio map. The ratio of each pixel represents the degree of difference in the spatial change rate of the two physical channels at that location.

[0074] In the defect edge region, the ratio increases significantly because the intensity of diffuse reflection light changes drastically and the intensity gradient of nonpolarized light is large, while the spatial variation of polarization modulation amplitude is relatively gentle and the gradient amplitude is small. In the flat and uniform texture region, the gradient amplitudes of the two channels are small and close in magnitude, and the ratio approaches 1. In the highly reflective residual region, the gradient amplitudes of the two channels may be large at the same time, and the ratio is close to 1, so no abnormally high weights will be generated.

[0075] The local gradient ratio map is normalized; the maximum and minimum values ​​of the ratios are counted by traversing the entire image, and the maximum-minimum normalization method is used to subtract the minimum value from the ratio of each pixel and divide by the difference between the maximum and minimum values ​​to map it to the value range of 0 to 1, thus obtaining the normalized ratio map.

[0076] The normalized ratio is combined with a preset weight mapping curve to generate an adaptive weight map. The weight mapping curve adopts the form of a piecewise linear function or an exponential saturation function: for pixels with a normalized ratio lower than the first threshold, the weight output is the baseline weight 1; for pixels with a normalized ratio between the first and second thresholds, the weight increases linearly or exponentially with the ratio; for pixels with a normalized ratio higher than the second threshold, the weight output is the preset maximum weight to prevent excessive local weight from causing over-enhancement and artifacts; the first and second thresholds are calibrated according to the protective layer material type and the contrast characteristics of typical defect edges, and the baseline weight 1 and the maximum weight range from 1 to 3.

[0077] The generated adaptive weight map is smoothed by Gaussian filtering, with the standard deviation of the filter kernel set to 1 to 2 pixels. The purpose of smoothing is to eliminate the discrete noise introduced by the pixel-by-pixel ratio calculation, so that the weights transition smoothly in the neighborhood of the defect edge and avoid blocky splicing marks in the fused image. The weight values ​​are not renormalized after smoothing, so as to maintain the original numerical magnitude relationship.

[0078] The smoothed adaptive weight map is output to the subsequent fusion enhancement step. The size of the adaptive weight map is exactly the same as that of the unpolarized light intensity component image and the polarization modulation amplitude image, and the pixel space corresponds one-to-one. The weight value of each pixel position is the weighting coefficient of the polarization modulation amplitude information at that position during fusion. Positions with a weight greater than 1 correspond to the defect edge enhancement region, and positions with a weight equal to or close to 1 correspond to the normal texture preservation region, thus realizing full-image adaptive detail enhancement control.

[0079] The process of generating defect-enhanced images is as follows: Three spatially registered image data of the same size are obtained from the multi-frame polarization fusion processing flow: the unpolarized light intensity component image, with pixel values ​​of floating-point grayscale, representing the pure diffuse light intensity base after removing specular reflection components; the polarization modulation amplitude image, with pixel values ​​of floating-point grayscale, representing the strength of polarization modulation components in the reflected light of each pixel; and the adaptive weight map, with pixel values ​​of floating-point weight coefficients, which have been magnified in the defect edge region and maintained as the baseline value in the flat texture region.

[0080] All three images have a row and column size of W columns × H rows, with pixels corresponding point by point and consistent coordinate indices.

[0081] The grayscale value of the polarization modulation amplitude image and the corresponding weight coefficient in the adaptive weight image are read pixel by pixel, and the two are multiplied to obtain the weighted polarization modulation amplitude image. The weighting operation amplifies the polarization modulation amplitude of high-weight areas such as defect edges and coating abrupt changes, while the polarization modulation amplitude of low-weight areas in flat texture areas is maintained or slightly attenuated. The weighting process does not change the spatial structure of the image, but only modulates the polarization detail contribution of each pixel.

[0082] The grayscale values ​​of the unpolarized light intensity component image are read pixel by pixel and added to the corresponding pixel grayscale values ​​of the weighted polarization modulation amplitude image. The superposition follows the fusion paradigm of the base layer plus the detail layer. The unpolarized light intensity component image provides a global diffuse reflection base, and the weighted polarization modulation amplitude image provides adaptively filtered polarization edge details.

[0083] Before superposition, the weighted polarization modulation amplitude image is multiplied by a global enhancement factor. This enhancement factor is an adjustable parameter, preset according to the type of protective layer, and is used to control the intensity of detail enhancement as a whole. The typical value range is 0.5 to 1.5. If the enhancement factor is too low, the defect edge enhancement will be insufficient, and if it is too high, overshoot artifacts may be introduced at strong edges.

[0084] The superimposed pixel grayscale values ​​may exceed the normal image display range. At the edge of the defect, the gradient is extremely large or the weight is high, the grayscale value may exceed the upper limit, and in the area of ​​extremely low reflectivity, the grayscale value may be lower than the lower limit.

[0085] The fusion result is dynamically cropped, with the lower limit of the output grayscale set to 0 and the upper limit determined by multiplying the global maximum grayscale value of the unpolarized light intensity component image by a preset expansion coefficient. The preset expansion coefficient controls the cropping tolerance, with a value range of 1.1 to 1.3. Pixels with grayscale values ​​less than the lower limit are set to the lower limit value, pixels with grayscale values ​​greater than the upper limit are set to the upper limit value, and pixels in between remain unchanged. The cropping operation preserves the enhancement effect while preventing secondary information loss caused by local overexposure.

[0086] Linear grayscale stretching is performed on the cropped and fused image to map the image grayscale values ​​from the current range to the standard output range. The stretching adopts the maximum-minimum linear mapping method. First, the minimum and maximum grayscale values ​​of the entire image after cropping are calculated. Then, the minimum value is subtracted from the grayscale value of each pixel, divided by the difference between the maximum and minimum values, multiplied by the target range width, and finally the lower limit of the target range is added. The target output range is set according to the input requirements of the subsequent defect detection network.

[0087] The floating-point grayscale values ​​are rounded to the nearest integer and converted to the integer pixel format required by the subsequent defect detection network. This completes the bit depth conversion and generates a standard format defect enhancement image.

[0088] The generated enhanced image of defects is output to the defect recognition module. The image has the following characteristics: the specular reflection highlight component is effectively removed, the original texture and color information of the protective layer are completely preserved by the non-polarized substrate, and the details of minor defects such as scratch edges, pinhole outlines, and uneven coating boundaries are significantly enhanced after weighted superposition of polarized details. The overall image contrast and defect recognition are significantly better than the original single-frame polarized image or the image that has only undergone orthogonal extinction processing.

[0089] The defect recognition module inputs the enhanced defect image into a pre-trained defect detection network and outputs the defect category, location, and confidence level. The process of outputting defect category, location, and confidence level is as follows: A defect detection model based on a deep convolutional neural network is constructed. The overall network adopts a cascaded architecture of backbone network, feature pyramid, and detection head.

[0090] The backbone network adopts a residual network structure, which contains multiple residual stages. Each stage is composed of several stacked residual blocks. The residual blocks are connected sequentially with convolutional layers, batch normalization layers and activation function layers. The input is directly passed to the output through skip connections to alleviate the gradient vanishing problem of deep networks. The four stages of the backbone network output feature maps with spatial resolution halved and channel number doubled step by step, forming a multi-scale feature hierarchy.

[0091] The feature pyramid network receives the output feature maps from the four stages of the backbone network and performs multi-scale feature fusion through a top-down upsampling path and lateral connections. After upsampling, the high-level semantic features are added element-wise with the adjacent low-level detail features to generate a set of scale-rich and semantically enhanced fused feature maps. Each layer of the fused feature map is responsible for the response to defects of different sizes, with low-level high-resolution features responsible for small-sized defects and high-level low-resolution features responsible for large-sized defects.

[0092] The detection head adopts a two-stage design of a region proposal network and a classification and regression sub-network. The region proposal network slides anchor boxes of preset size and aspect ratio on the fused feature maps of each layer to determine whether the anchor boxes contain defects and initially regress the bounding box offset to filter out candidate proposal regions. The classification and regression sub-network uses the region of interest alignment operation to map the candidate proposal regions onto the corresponding layer feature maps to extract fixed-size feature vectors. After passing through a fully connected layer, it outputs two sets of predictions in parallel: one set is the probability distribution of defect categories, and the other set is the fine regression parameters of the bounding boxes.

[0093] Collect samples of various typical defects in the protective layer of aluminum alloy flooring, including common defect types such as scratches, bubbles, orange peel, pinholes, coating peeling, and uneven coating; each defect sample is processed by the system's pre-processing module to generate a corresponding defect enhancement image, which constitutes the training image set.

[0094] Each training image is manually annotated in detail. The annotation information includes: the defect category name, coded with a preset category number; the minimum bounding rectangle of the defect region, described by four parameters: the horizontal and vertical coordinates of the top left corner of the rectangle, the width of the rectangle, and the height of the rectangle; after annotation, a training dataset is formed with image files and annotation files corresponding one-to-one, and is randomly divided into training set and validation set according to a preset ratio.

[0095] The pre-trained weights of the backbone network on a publicly available large-scale image classification dataset are used as initialization parameters, while the parameters of the remaining network layers are randomly initialized. The training adopts the mini-batch stochastic gradient descent optimization algorithm, and the loss function is composed of the foreground-background binary classification loss and bounding box regression loss of the region proposal network, as well as the multi-class cross-entropy loss and bounding box regression loss of the classification regression sub-network.

[0096] In each training round, all samples in the training set are traversed in batches. Forward propagation calculates the loss between the predicted output and the labeled ground truth, and backpropagation calculates the gradient and updates the network weights. After each training round, the model performance metrics are evaluated on the validation set, and the average precision and recall for various defects are recorded. When the overall metrics on the validation set no longer improve for several consecutive rounds, the early stopping mechanism is triggered to terminate the training, and the best-performing model weight file is saved as the pre-trained model.

[0097] To improve the model's generalization ability, online data augmentation is performed on the input images during training, including random horizontal flipping, random rotation by a preset angle, random brightness and contrast fine-tuning, and random scaling. The augmented images are then fed into the network for training in real time.

[0098] After the system enters the online detection stage, the defect enhancement image output by the multi-frame polarization fusion and defect enhancement module is used as the input to be detected. First, the pixel values ​​of the enhancement image are scaled and offset according to the normalization parameters during training to make the input data distribution consistent with the training stage. Then, the image spatial size is adjusted to the fixed size required by the network input layer, and the scaling method that keeps the aspect ratio unchanged is combined with edge filling.

[0099] The preprocessed image is fed into a defect detection network with pre-trained weights and performs a single forward inference computation. The backbone network extracts image features layer by layer, the feature pyramid performs multi-scale fusion, the region proposal network generates candidate defect regions on the feature maps of each layer and filters low-confidence proposals, and the classification and regression sub-network performs category determination and bounding box refinement on the retained candidate regions. The entire computation process is completed on a graphics processor, and the inference time per frame meets the real-time requirements.

[0100] The forward inference outputs a large number of candidate bounding boxes and their corresponding class probabilities and regression offset parameters. First, the bounding box regression parameters are applied to the corresponding anchor boxes to obtain the precise coordinates of the predicted bounding boxes. Then, non-maximum suppression is performed to remove duplicate bounding boxes with a spatial overlap ratio exceeding a preset threshold within the same class, retaining the box with the highest confidence.

[0101] The detection results after non-maximum suppression are filtered by a confidence threshold. A defect judgment confidence threshold is set, typically 0.5 to 0.7. Only detection results with a class probability higher than the threshold are retained as valid defect outputs, and unreliable detections with too low confidence are removed.

[0102] Each valid defect detection result after filtering is packaged and output in a fixed data format. Each result contains three fields: defect category, identified by category name or number; defect location, described by the coordinates of the upper left and lower right corners of the bounding box; and confidence level, a floating-point number between 0 and 1 representing the degree of certainty of the model regarding the detection result.

[0103] All valid defect detection results are summarized into a defect list for the current frame and sent to the host computer display interface for visual overlay marking. At the same time, an audible and visual alarm is triggered and the defect image is archived, completing the complete detection process from image acquisition to defect recognition output.

[0104] Example 2: Please see Figure 2 Based on Example 1, Example 2 of this application also provides a method for image recognition of defects in the protective layer of aluminum alloy flooring, including the following specific steps: Step 1: Obtain the 3D point cloud data of the aluminum alloy floor protective layer surface, calculate the local normal vector and the high reflectivity risk coefficient, and divide the field of view into a low curvature region and a high curvature region to be scanned based on the high reflectivity risk coefficient. Step 2: Calculate the optimal extinction angle for the low curvature region based on the local normal vector and fix the polarization angle to acquire a single-frame polarization image. Perform a gradual polarization angle scan on the high curvature region to be scanned with the optimal extinction angle as the center, and acquire a multi-frame polarization image sequence. Step 3: Establish a physical model of light intensity variation with polarization angle for each pixel of the multi-frame polarization image sequence. Separate the non-polarized light intensity component image and the polarization modulation amplitude image by fitting. Use the single-frame polarization image as the non-polarized light intensity component in the low curvature region and synthesize it with the non-polarized light intensity component image in the high curvature region to form a full-field non-polarized base image. Construct adaptive weights using the dual-channel local gradient ratio and weightedly fuse the polarization modulation amplitude image into the full-field non-polarized base image to generate a defect enhancement image. Step 4: Input the enhanced image of the defect into the pre-trained defect detection network, and output the defect category, location, and confidence level.

[0105] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented in software, the above embodiments can be implemented, in whole or in part, as a computer program product. Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution.

[0106] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0107] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A defect image recognition system for aluminum alloy floor protective layers, characterized in that, The system includes: The 3D shape pre-sensing module acquires 3D point cloud data of the surface of the aluminum alloy floor protective layer, calculates the local normal vector and high reflectivity risk coefficient, and divides the field of view into a low curvature region and a high curvature region to be scanned based on the high reflectivity risk coefficient. The adaptive polarization control imaging module calculates the optimal extinction angle for low curvature regions based on local normal vectors and acquires a single-frame polarization image with a fixed polarization angle. For high curvature regions to be scanned, it performs a gradual polarization angle scan centered on the optimal extinction angle and acquires a multi-frame polarization image sequence. The multi-frame polarization fusion and defect enhancement module establishes a physical model of light intensity variation with the detection angle for each pixel of the multi-frame polarization image sequence. By fitting, it separates the non-polarized light intensity component image and the polarization modulation amplitude image. The single-frame polarization image is used as the non-polarized light intensity component in the low curvature region and synthesized with the non-polarized light intensity component image in the high curvature region to form a full-field non-polarized base image. The module uses the local gradient ratio of the two channels to construct adaptive weights and weightedly fuses the polarization modulation amplitude image into the full-field non-polarized base image to generate a defect enhancement image. The defect recognition module inputs the enhanced defect image into a pre-trained defect detection network and outputs the defect category, location, and confidence level.

2. The image recognition system for defects in the protective layer of aluminum alloy flooring according to claim 1, characterized in that, The process of calculating the local normal vector and the high reflectivity risk coefficient is as follows: Based on pixel-by-pixel analysis of 3D point clouds, depth gradients are derived. Normal vectors are constructed using negative horizontal and vertical gradient values ​​and normalized. After mean filtering, a smoothed normal vector field is obtained. Based on the smoothed normal vector field, the rate of change of each component of the normal vector with respect to the horizontal and vertical coordinates is obtained. The sum of the horizontal and vertical moduli of change is taken as the high reflectivity risk coefficient. The high reflectivity risk coefficient is compared with a preset threshold to divide the low curvature region and the high curvature region to be scanned.

3. The image recognition system for defects in the protective layer of aluminum alloy flooring according to claim 2, characterized in that, The process of dividing the field of view into a low-curvature region and a high-curvature region to be scanned is as follows: The high reflectivity risk coefficient of each pixel is compared with a preset threshold to determine high curvature pixels and low curvature pixels; Connectivity analysis is performed on high curvature pixels. After removing isolated connected components with an area smaller than a preset threshold, a binary region label map is generated. High curvature regions are assigned a value of 1, and low curvature regions are assigned a value of 0.

4. The image recognition system for defects in the protective layer of aluminum alloy flooring according to claim 3, characterized in that, The process of calculating the optimal extinction angle for low curvature regions based on local normal vectors and acquiring a single-frame polarization image with a fixed analyzer angle is as follows: Read the low curvature region marked with a value of 0 in the region marker map and extract the smooth normal vector within the region; calculate the average value of the vertical deflection angle of the normal vector within the region as the equivalent tilt angle, calculate the average value of the arctangent of the ratio of the vertical component to the horizontal component of the normal vector as the equivalent polarization rotation angle, and superimpose the base extinction angle with the equivalent polarization rotation angle to obtain the optimal extinction angle for the region; adjust the analyzer to the optimal extinction angle and fix the starting angle, and trigger the camera to acquire a single frame polarization image.

5. The image recognition system for defects in the protective layer of aluminum alloy flooring according to claim 4, characterized in that, The process of performing a gradient scan of the high-curvature scanning area with the optimal extinction angle as the center is as follows: The median of the optimal extinction angle distribution of each pixel in the high curvature connected region is taken as the reference optimal extinction angle; the scanning range is extended to both sides by a preset multiple standard deviation with the reference optimal extinction angle as the center, and the scanning step size is determined according to the polarization extinction half-width angle; an equal arithmetic polarization detection angle scanning sequence is generated with the reference optimal extinction angle as the center; the starting angle is fixed, the polarization detection angle is adjusted frame by frame according to the scanning sequence and polarization images are acquired, and the multi-frame image sequence and the corresponding polarization detection angle sequence are packaged and output.

6. The image recognition system for defects in the protective layer of aluminum alloy flooring according to claim 1, characterized in that, The process of separating the unpolarized light intensity component image and the polarization modulation amplitude image by fitting is as follows: N frames of light intensity vectors and corresponding polarization detection vectors are extracted pixel by pixel; the light intensity model is linearly expanded into a combination of three basis functions, and an N×3 design matrix is ​​constructed. The three combination coefficients are solved at once by multiplying the pre-stored inverse matrix with the light intensity vector; the unpolarized light intensity component is obtained by subtracting the polarization modulation amplitude from the first combination coefficient, and the square root of the sum of the squares of the second and third combination coefficients is taken as the polarization modulation amplitude; the fitting residual between the measured light intensity value and the model fitting value is calculated pixel by pixel. When the fitting residual exceeds the preset residual tolerance threshold, the polarization response of the pixel is determined to deviate from the ideal model. The minimum value of the light intensity sequence is taken as the unpolarized light intensity of the pixel, and the polarization modulation amplitude is set to the preset minimum value.

7. The image recognition system for defects in the protective layer of aluminum alloy flooring according to claim 6, characterized in that, The process of generating defect-enhanced images is as follows: Calculate the local gradient magnitude of the unpolarized light intensity component image and the polarization modulation amplitude image, calculate the ratio of their gradient magnitudes pixel by pixel and normalize it to construct an adaptive weight map; multiply the polarization modulation amplitude image pixel by pixel by the adaptive weight map and the preset global enhancement coefficient, and then superimpose it pixel by pixel with the unpolarized light intensity component image. After dynamic range cropping and grayscale stretching, a defect enhancement image is generated.

8. The image recognition system for defects in the protective layer of aluminum alloy flooring according to claim 1, characterized in that, The process of outputting defect category, location, and confidence level is as follows: A defect detection network is adopted, which consists of a backbone network, a feature pyramid, and a detection head cascaded together. The enhanced defect image is input into the network, and after multi-scale feature extraction and fusion, the region proposal network generates candidate regions, and the classification and regression sub-network outputs the defect category probability and bounding box parameters. After nonmaximum suppression and confidence threshold filtering, the defect category, location coordinates, and confidence level are output.

9. A method for image recognition of defects in the protective layer of aluminum alloy flooring, applied to the image recognition system for defects in the protective layer of aluminum alloy flooring as described in any one of claims 1-8, characterized in that, The method includes: Step 1: Obtaining three-dimensional point cloud data of the surface of the aluminum alloy floor protective layer, calculating the local normal vector and the high reflectivity risk coefficient, and dividing the field of view into a low curvature region and a high curvature region to be scanned based on the high reflectivity risk coefficient; Step 2: Calculate the optimal extinction angle for the low curvature region based on the local normal vector and fix the polarization angle to acquire a single-frame polarization image. Perform a gradual polarization angle scan on the high curvature region to be scanned with the optimal extinction angle as the center, and acquire a multi-frame polarization image sequence. Step 3: Establish a physical model of light intensity variation with polarization angle for each pixel of the multi-frame polarization image sequence. Separate the non-polarized light intensity component image and the polarization modulation amplitude image by fitting. Use the single-frame polarization image as the non-polarized light intensity component in the low curvature region and synthesize it with the non-polarized light intensity component image in the high curvature region to form a full-field non-polarized base image. Construct adaptive weights using the dual-channel local gradient ratio and weightedly fuse the polarization modulation amplitude image into the full-field non-polarized base image to generate a defect enhancement image. Step 4: Input the enhanced image of the defect into the pre-trained defect detection network, and output the defect category, location, and confidence level.