A lug adapter sheet surface defect detection method based on deep learning

By employing multi-view polarization imaging and photometric distribution model decoupling techniques, the problems of surface morphology distortion and photometric interference in the detection of surface defects of automotive tab adapters were solved, achieving high-precision defect feature identification.

CN122289274APending Publication Date: 2026-06-26GANGYANG AXIAN TECH (GUANYUN) CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GANGYANG AXIAN TECH (GUANYUN) CO LTD
Filing Date
2026-05-29
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

When inspecting automotive electrode adapters, existing technologies struggle to achieve accurate decoupling and spatially consistent mapping of the surface features of the target under conditions of curved surface distortion and complex photometric interference, leading to difficulties in identifying defect features.

Method used

Three-dimensional geometric curvature parameters are obtained through multi-view polarization imaging. The specular reflection and diffuse structure features are decoupled using a photometric distribution model. Geometric distortion compensation and light and shadow artifact filtering are performed by combining the local tangent space mapping matrix and the anisotropic structure tensor. The results are then input into a depth recognition network for feature extraction and classification.

Benefits of technology

It effectively separates the specular highlight area from the microscopic morphological features, compensates for image distortion, ensures the physical reliability and consistency of the features, and improves the accuracy of defect detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122289274A_ABST
    Figure CN122289274A_ABST
Patent Text Reader

Abstract

This invention relates to the field of image analysis technology, specifically a deep learning-based method for detecting surface defects in tab adapters. The method includes acquiring the three-dimensional geometric curvature parameters of the bending region through multi-view polarization imaging; decoupling specular reflection and diffuse structural features using a photometric distribution model to resolve high-brightness reflection interference; nonlinearly projecting the surface texture onto a two-dimensional plane using a local tangent space mapping matrix to generate an isotropic corrected texture map to compensate for geometric distortion; identifying suspected defect points that disrupt flow continuity through singular value decomposition of the anisotropic structure tensor; and introducing stress distribution priors and epipolar geometric constraints for multi-view spatial consistency verification to effectively filter lighting artifacts. The verified region and corrected texture are fused and input into a deep recognition network with a spatial attention mechanism to analyze the surface defect image. This invention achieves accurate analysis of surface defect images through deep coupling of three-dimensional geometry and a photometric physical model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image analysis technology, specifically to a method for detecting surface defects in tab adapters based on deep learning. Background Technology

[0002] With the widespread application of computer vision technology in industrial automation inspection, using deep learning models to identify surface defects in metal parts has become a mainstream technical direction. When inspecting automotive metal parts with complex geometries (such as bends and arc surfaces) and highly reflective surfaces (such as tab adapters), image analysis algorithms face severe challenges in feature representation.

[0003] In conventional image processing, due to the nonlinear superposition of specular and diffuse reflection components on metal surfaces, the acquired original image sequences often suffer from severe brightness saturation and texture detail occlusion. Furthermore, when the object to be detected has significant three-dimensional curvature changes, the receptive field of traditional two-dimensional convolution operators in non-Euclidean space will be distorted, making it difficult to maintain the spatial consistency of defect features during projection transformation. In addition, due to the natural texture interference caused by the flow direction of metal grains and the non-uniform distribution of the ambient light field, deep learning models often struggle to accurately distinguish between real geometric damage and light and shadow artifacts from complex background noise when extracting weak defect features.

[0004] In summary, how to achieve accurate decoupling and spatially consistent mapping of the surface features of the target under the conditions of curved morphology distortion and complex photometric interference is a technical problem that urgently needs to be solved in the field of computer vision.

[0005] To address this, a deep learning-based method for detecting surface defects in tab adapters is proposed. Summary of the Invention

[0006] The purpose of this invention is to provide a deep learning-based method for detecting surface defects in tab adapters. By deeply coupling three-dimensional geometry and a photometric physical model, accurate analysis of curved surface defect images is achieved. This includes obtaining the three-dimensional geometric curvature parameters of the bending region through multi-view polarization imaging; decoupling specular reflection and diffuse structural features using a photometric distribution model to resolve high-brightness reflection interference; nonlinearly projecting the curved surface texture onto a two-dimensional plane using a local tangent space mapping matrix to generate an isotropic corrected texture map to compensate for geometric distortion; identifying suspected defect points that disrupt flow continuity through singular value decomposition of the anisotropic structure tensor; and introducing stress distribution priors and epipolar geometric constraints for multi-view spatial consistency verification to effectively filter lighting artifacts. The verified region and corrected texture are fused and input into a deep recognition network with a spatial attention mechanism to analyze the surface defect image.

[0007] To achieve the above objectives, the present invention provides the following technical solution: A deep learning-based method for detecting surface defects in electrode adapters, comprising: Obtain multi-view image sequences of the bending area of ​​the electrode adapter piece and the corresponding three-dimensional geometric curvature parameters; Calculate the pixel brightness variance of the multi-view image sequence at the same spatial coordinates, establish a photometric distribution model, and separate the specular reflection component field and the diffuse structure feature field from the multi-view image sequence; A local tangent space mapping matrix is ​​constructed using the three-dimensional geometric curvature parameters, and the diffuse structural feature field is projected from the surface space to the two-dimensional Euclidean plane to generate a corrected texture map. The second-order partial derivative of the corrected texture map is calculated and an anisotropic structure tensor is constructed. The gradient field reflecting the flow direction of metal grains is extracted, and the candidate pixel points of defects are determined according to the singular values ​​of the gradient field. Based on the three-dimensional geometric curvature parameters, a stress distribution prior mask and epipolar geometric space constraint criteria are constructed. The prior mask and the constraint criteria are used to perform multi-view spatial consistency verification on the candidate defect pixels. The verified pixel regions are fused with the corrected texture map and input into a pre-trained deep defect recognition network for feature extraction and classification, and the surface defect detection results are output.

[0008] Preferably, the multi-view image sequence is acquired by an equidistant ring array deployed above the bending area of ​​the tab adapter, including: The ridge-line frontal view, perpendicular to the central normal of the arc-shaped ridge in the bending region, is used to capture the strong reflected brightness field at the bend apex; the flank symmetrical view, an oblique observation angle symmetrically distributed on both sides of the arc-shaped ridge, is used to obtain the microscopic grain flow characteristics of the bend transition slope; the root grazing view, a low-elevation observation position located at the junction of the bending region and the flat substrate, highlights the microcracks and indentations caused by stress concentration through the shadow enhancement effect formed by grazing light.

[0009] Preferably, the process of acquiring the multi-view image sequence and the corresponding three-dimensional geometric curvature parameters includes: simultaneously acquiring sub-images of the tab adapter in multiple orthogonal polarization states by configuring polarization modulation elements at the front end of the camera device at each observation position; each frame of the multi-view image is processed by multi-channel fusion to generate a composite physical feature map containing polarization degree and phase delay, which is used to distinguish the specular bright spot area from the actual metal surface micro-damage by utilizing the optical anisotropy of the metal surface under different bending curvatures; constructing a rough depth manifold of the bending region by utilizing the parallax relationship between the symmetrical viewpoints of the flanks, and inverting the surface normal vector by utilizing the polarization component under the frontal viewpoint of the ridge to complete the depth manifold; calculating the Gaussian curvature and average curvature at each pixel coordinate point by performing differential geometric operations on the reconstructed three-dimensional surface, as the three-dimensional geometric curvature parameters.

[0010] Preferably, the process of establishing the photometric distribution model includes: extracting pixel brightness vectors at the same spatial coordinates in the multi-view image sequence, calculating the pixel brightness variance at each coordinate point to obtain statistical parameters characterizing the dispersion of the brightness distribution; calculating the local normal change rate of each pixel in the bending region using the three-dimensional geometric curvature parameters, and using the normal change rate as a constraint operator to perform spatial weight correction on the statistical variance; establishing a brightness response probability model about the viewpoint vector based on the corrected statistical variance, and generating an expected envelope surface describing the ideal brightness distribution of each pixel position under different observation orientations, as the photometric distribution model.

[0011] Preferably, the process of separating the specular reflection component field and the diffuse structure feature field from the multi-view image sequence includes: calculating the photometric response residual of the measured pixel brightness in the multi-view image sequence relative to the expected envelope surface in the photometric distribution model; constructing an anisotropic threshold function in combination with the three-dimensional geometric curvature parameters, performing logical judgment on the photometric response residual, and extracting outlier pixels exceeding the threshold as the specular reflection component field reflecting the instantaneous strong light on the metal surface; stripping the specular reflection component field from the original multi-view image sequence, and performing cross-view consistency feature enhancement on the remaining pixel components, outputting a diffuse structure feature field characterizing the micro-texture and diffuse reflection structure of the metal surface.

[0012] Preferably, the process of generating the corrected texture map includes: determining the differential manifold structure of each sampling point in the diffuse structure feature field using the three-dimensional geometric curvature parameters, and calculating the local metric tensor characterizing the degree of surface stretching or compression; using the geometric center line of the bending region as a reference, calculating the projection transformation coefficients of each pixel point on the tangent plane using the local metric tensor, and thereby constructing the local tangent space mapping matrix of each pixel position; using the local tangent space mapping matrix to establish the nonlinear mapping relationship between the three-dimensional surface coordinates and the two-dimensional Euclidean plane coordinates of each sampling point in the diffuse structure feature field, and projecting the surface texture onto the standard two-dimensional coordinate system; establishing an orthogonal pixel grid in the two-dimensional Euclidean plane, resampling the projected feature components using a bilinear interpolation algorithm, filling the image pixel holes caused by the abrupt curvature change, and outputting the isotropic corrected texture map.

[0013] Preferably, the process of determining candidate defect pixels based on the singular values ​​of the gradient field includes: performing multi-scale Gaussian convolution on the corrected texture map, calculating the second-order partial derivatives of each pixel coordinate point in the horizontal, vertical, and diagonal directions, and constructing a Hessian matrix representing the local gray-level change rate; constructing a local structure tensor using the elements of the Hessian matrix, and performing spatial smoothing integration on the structure tensor using a preset window to fuse the metal grain flow direction features of adjacent regions; performing singular value decomposition on the integrated structure tensor, extracting the principal singular values ​​reflecting the texture energy intensity and the feature vectors reflecting the consistency of the texture direction, and constructing a gradient field; calculating the anisotropy confidence of each pixel position in the gradient field, and determining the pixel positions where the confidence is lower than a preset threshold and whose principal singular values ​​undergo abrupt changes as candidate defect pixels, thereby identifying local morphological anomalies that disrupt the continuity of the metal grain flow direction.

[0014] Preferably, the multi-view spatial consistency verification includes: mapping the Gaussian curvature and mean curvature in the three-dimensional geometric curvature parameters to physical stress concentration probabilities, generating the stress distribution prior mask; using the relative pose matrix of the camera device when acquiring multi-view image sequences, calculating the epipolar equation between each observation viewpoint, and using it as the epipolar geometric spatial constraint criterion; performing a one-dimensional search along the epipolar line in the images of adjacent viewpoints, calculating the cross-view feature matching degree of the defect candidate pixels, eliminating pseudo-defect points that cannot satisfy the geometric projection constraints in multiple viewpoints or whose comprehensive score after combining confidence weights is lower than a set threshold, and retaining the real defect pixel region with multi-view spatial consistency.

[0015] Preferably, the output surface defect detection result includes: using the real defect pixel region as a spatial guiding channel, and splicing and fusing it with the corrected texture map containing local topological information in the channel dimension to construct a multi-channel feature input tensor; inputting the multi-channel feature input tensor into a pre-trained deep defect recognition network, the deep defect recognition network including a backbone network, a spatial attention injection module, a multi-scale feature pyramid, and a classification and regression head network; the backbone network includes a first feature extraction stage, a second feature extraction stage, a third feature extraction stage, and a fourth feature extraction stage connected in sequence; the first feature extraction stage extracts the multi-channel feature input tensor and outputs a first feature map; the spatial attention injection module scales the spatial size of the stress distribution prior mask to match the first feature map. Figure 1 The first feature map is expanded along the channel dimension and then multiplied element-wise with the second feature map to generate a weighted feature map with attention weights. The second feature extraction stage receives the weighted feature map and extracts features, outputting a second feature map. The third feature extraction stage receives the second feature map and extracts features, outputting a third feature map. The fourth feature extraction stage receives the third feature map and extracts features, outputting a fourth feature map. The multi-scale feature pyramid receives the second, third, and fourth feature maps, and fuses them to construct a multi-scale pyramid feature map through convolutional dimensionality reduction and top-down upsampling and element-wise addition. The multi-scale pyramid feature map is input into the classification and regression head network, which outputs specific category labels and localization bounding boxes representing microcracks, indentations, or scratches on the surface of the tab adapter.

[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By introducing polarization modulation and phase delay features into the multi-view acquisition sequence, and utilizing the mapping relationship between optical anisotropy and curvature of the metal surface, the effective separation of the specular highlight area and micro-morphological features at the physical property level was achieved. Combined with the inversion of the flank parallax manifold and the frontal normal vector, dense completion of surface geometric parameters was completed in three-dimensional space. This provides high-fidelity curvature a priori information for subsequent algorithms, enabling the accurate digital reconstruction of the three-dimensional topological features of the surface under test, thus laying the physical information foundation for multi-dimensional feature fusion.

[0017] 2. By constructing a tangent space mapping matrix using local metric tensors, the diffuse structural feature field on a 3D curved manifold is nonlinearly projected onto a 2D Euclidean plane, compensating for radial and axial distortions in the image caused by part bending. Pixel holes in regions of abrupt curvature changes are eliminated through resampling of the forward pixel grid and bilinear interpolation, achieving a standardized transformation from non-Euclidean space texture to isotropic corrected texture. This ensures the topological invariance of local features before and after the projection transformation, enabling subsequent convolution operators to extract morphologically stable geometric features within a unified coordinate system.

[0018] 3. Macroscopic geometric curvature is transformed into microscopic stress distribution probability, and the relative pose of the camera device is coupled to construct epipolar geometric space constraints, establishing a cross-view feature mapping association. Through one-dimensional search along the epipolar line and confidence weight determination, geometric closed-loop verification of defect candidate points in multi-dimensional projection space is achieved. From the data source, instantaneous reflection noise, dust artifacts, and real structural damage are logically distinguished, effectively filtering out false responses that do not conform to the projection topology, strengthening the consistency between the data to be processed and the physical damage model, and ensuring that the features entering the neural network have extremely high physical credibility. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of a deep learning-based method for detecting surface defects in a tab adapter plate according to the present invention. Figure 2 This is a schematic diagram of the process for separating the mirror reflection component field and the diffuse structure feature field of the present invention; Figure 3 This is a schematic diagram of the multi-view spatial consistency verification process of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] Please see Figures 1 to 3 This invention provides a deep learning-based method for detecting surface defects in electrode adapters, the technical solution of which is as follows:

[0022] Example 1: A deep learning-based method for detecting surface defects in electrode adapters, comprising: Obtain multi-view image sequences of the bending area of ​​the electrode adapter piece and the corresponding three-dimensional geometric curvature parameters; Calculate the pixel brightness variance of the multi-view image sequence at the same spatial coordinates, establish a photometric distribution model, and separate the specular reflection component field and the diffuse structure feature field from the multi-view image sequence; A local tangent space mapping matrix is ​​constructed using the three-dimensional geometric curvature parameters, and the diffuse structural feature field is projected from the surface space to the two-dimensional Euclidean plane to generate a corrected texture map. The second-order partial derivative of the corrected texture map is calculated and an anisotropic structure tensor is constructed. The gradient field reflecting the flow direction of metal grains is extracted, and the candidate pixel points of defects are determined according to the singular values ​​of the gradient field. Based on the three-dimensional geometric curvature parameters, a stress distribution prior mask and epipolar geometric space constraint criteria are constructed. The prior mask and the constraint criteria are used to perform multi-view spatial consistency verification on the candidate defect pixels. The verified pixel regions are fused with the corrected texture map and input into a pre-trained deep defect recognition network for feature extraction and classification, and the surface defect detection results are output.

[0023] Furthermore, the multi-view image sequence is acquired through a ring-shaped equidistant array deployed above the bending area of ​​the tab adapter, including: The ridge-line frontal view, perpendicular to the central normal of the arc-shaped ridge in the bending region, is used to capture the strong reflected brightness field at the bend apex; the flank symmetrical view, an oblique observation angle symmetrically distributed on both sides of the arc-shaped ridge, is used to obtain the microscopic grain flow characteristics of the bend transition slope; the root grazing view, a low-elevation observation position located at the junction of the bending region and the flat substrate, highlights the microcracks and indentations caused by stress concentration through the shadow enhancement effect formed by grazing light.

[0024] First, a hemispherical or annular mechanical support frame is constructed around the bending center axis of the tab adapter. On this support frame, at least five industrial camera units are deployed, with the arc-shaped ridge of the bending area as the geometric origin. The first camera unit is fixed on the central normal directly above the bending apex, forming a frontal view of the ridge. The second and third cameras are installed symmetrically on either side at angles of 30 to 60 degrees relative to the central normal, forming a symmetrical flank view. The fourth and fifth cameras are installed at low pitch angles close to the flat substrate surface, with the pitch angle controlled between 10 and 20 degrees, forming a root-grazing view. The optical axes of all cameras point to the geometric center of the bending area, ensuring that the multi-view observation vectors intersect at the origin of the spatial coordinate system.

[0025] A four-way beam polarization sensor (or an electronically controlled liquid crystal polarization rotator installed at the front of the lens) is configured at the front of each viewing angle camera device, and combined with a ring-shaped shadowless light source or a strip-shaped combined light source. In order to achieve millisecond-level synchronous acquisition of multi-view image sequences, all cameras are connected to an FPGA-based industrial synchronization controller to ensure that the delay deviation of the trigger pulse between each camera is controlled within 10 microseconds. During the system initialization phase, the internal and external parameters of each viewing angle are jointly calibrated using a calibration board to determine the absolute coordinates C(x, y, z) of the optical center of each camera in the world coordinate system. This allows the acquisition of the rotation matrix and translation vector of each camera relative to the coordinate system of the center of the bending area. By adjusting the exposure time and gain of each viewing angle, pixel overflow is prevented in the strong reflection area under the frontal view of the ridge, while ensuring that the dark details under the grazing view at the root have sufficient contrast.

[0026] The strong reflective brightness field refers to the spatial distribution data of pixel grayscale values, characterized by the deformation features of the bending apex, formed by the dominant specular reflection component of the metal surface under the frontal view of the ridge. Specifically, the acquisition logic of this field is as follows: using a camera device deployed in the central normal direction, the peak-shaped brightness distribution at the ridge of the bending region due to high reflectivity is recorded under preset exposure time and gain parameters. This distribution data is stored in the form of a two-dimensional pixel matrix, where the value of each coordinate point represents the light intensity response intensity of that physical location under specular reflection conditions. This value is used as the reference input for the photometric distribution model in subsequent steps, thereby achieving accurate stripping of the specular reflection component.

[0027] In the actual acquisition process, when the tab adapter moves to the detection station and triggers the positioning sensor, the synchronous controller sends a hardware trigger pulse to all cameras in the array. The ridge-facing camera captures the brightness distribution information of the bending apex; the flank symmetrical camera group observes through bidirectional tilt to obtain the grain micro-texture of the bending slope after being stretched; and the root grazing camera uses low-angle grazing light to form a significant shadow enhancement effect at the tiny concave edge at the junction of the bending root and the substrate.

[0028] By constructing a ring-shaped, equidistant, multi-view acquisition array, layered observation of specific physical properties of the bending region was achieved, obtaining multi-dimensional raw information covering strong reflection fields, microscopic grain textures, and root shadow features. This viewpoint layout provides a complete spatial observation benchmark for subsequent photometric modeling and feature decoupling, ensuring the integrity of feature extraction required for underlying image analysis.

[0029] Furthermore, the process of acquiring the multi-view image sequence and the corresponding three-dimensional geometric curvature parameters includes: simultaneously acquiring sub-images of the tab adapter in multiple orthogonal polarization states by configuring polarization modulation elements at the front end of the camera device at each observation position; each frame of the multi-view image is processed by multi-channel fusion to generate a composite physical feature map containing polarization degree and phase delay, which is used to distinguish the specular bright spot area from the actual metal surface micro-damage by utilizing the optical anisotropy of the metal surface under different bending curvatures; constructing a rough depth manifold of the bending region by utilizing the parallax relationship between the symmetrical viewpoints of the flanks, and inverting the surface normal vector by utilizing the polarization component under the frontal viewpoint of the ridge to complete the depth manifold; calculating the Gaussian curvature and average curvature at each pixel coordinate point by performing differential geometric operations on the reconstructed three-dimensional surface, as the three-dimensional geometric curvature parameters.

[0030] While acquiring multi-view image sequences, the system simultaneously executes a three-dimensional geometric curvature parameter extraction process. Specifically, the front end of the camera device at each observation position is equipped with a four-directional beam-splitting polarization element, which can simultaneously capture sub-images of the electrode adapter in four orthogonal polarization states of 0 degrees, 45 degrees, 90 degrees, and 135 degrees in a single exposure. After receiving each group of sub-images, the back-end processing unit performs multi-channel fusion processing and uses Stokes vectors to calculate the polarization degree and phase delay information of each pixel position, thereby generating a composite physical feature map. This map utilizes the different curvatures of the metal surface. By analyzing the polarization anisotropy of incident light under varying polarization states, the system compares and analyzes the polarization angle shift between the specular reflection region and the background region. This allows for the precise differentiation between specular bright spots caused by metal bending and actual metal surface micro-damage caused by mechanical scratches, filtering out optical artifacts. Specifically, specular bright spots typically maintain a high degree of polarization, and their polarization angle changes smoothly with the surface normal. In contrast, actual metal micro-damage (such as scratches and indentations) experiences a drastic decrease in polarization (depolarization) due to microscopic diffuse reflection and multiple scattering, and the polarization angle exhibits random perturbations or abrupt changes. By setting a polarization gradient threshold and a local variance threshold for the polarization angle, high-polarization regions that conform to a smooth change pattern are identified as specular bright spots and removed from the background.

[0031] To construct a high-precision 3D topography of the bending region, the system first utilizes the parallax relationship between the symmetrical viewpoints of the left and right wings to calculate an initial depth map using a semi-global matching algorithm, thus constructing a coarse depth manifold for the bending region. Based on this, the system uses the polarization components obtained from the ridge's frontal viewpoint and combines them with a Fresnel reflection polarization model to invert the normal vector field of the metal surface. Since the depth manifold obtained by the parallax method is insensitive to minor surface undulations (such as microcracks), while the polarized normal vector is extremely sensitive to surface micro-slopes, the system employs a Poisson reconstruction algorithm to globally fuse high-frequency normal information with the low-frequency depth manifold, filling in and sharpening missing details in the depth manifold. Before performing the fusion, the parallax maps of the left and right wing views are back-projected onto the world coordinate system using pre-calibrated camera intrinsic and extrinsic matrices to generate a 3D point cloud. Simultaneously, based on the calibrated light source vector direction, the local polarized normal vectors at each pixel are transformed to a world coordinate system consistent with the point cloud through coordinate rotation transformation. By constructing a global energy functional, using the depth map as a position constraint and the polarization normal as a gradient constraint, the Laplace equation is solved to achieve seamless integration of the two.

[0032] After reconstructing a continuous 3D surface with subpixel-level accuracy, the system performs differential geometric operations on the 3D surface manifold. By calculating the coefficients of the first and second fundamental forms at each pixel coordinate point on the surface, the Gaussian curvature and mean curvature of that point are obtained and used as the 3D geometric curvature parameters. In specific implementation, the reconstructed 3D surface is first fitted locally (e.g., using a 3×3 or 5×5 window for bivariate quadratic polynomial fitting) to eliminate discrete sampling noise. Then, by obtaining the first and second partial derivatives of the fitted surface equation, the discretized first fundamental form parameters (E, F, G) and second fundamental form parameters (L, M, N) are obtained. Then, the Gaussian curvature K and mean curvature H of each point are calculated using the determinant ratio. These parameters can quantitatively characterize the degree of local deformation of the tab adapter during the bending process, providing a direct geometric basis for the subsequent construction of a priori stress distribution masks.

[0033] By spatially fusing polarization components and parallax depth, this application achieves physical identification of metal surface features and specular light and shadow, and uses differential geometric operators to characterize local deformation laws, establishing a topologically consistent description system, removing reflective interference, and providing objective quantitative basis for subsequent construction of stress distribution priors.

[0034] Furthermore, the process of establishing the photometric distribution model includes: extracting pixel brightness vectors at the same spatial coordinates in the multi-view image sequence, calculating the pixel brightness variance at each coordinate point to obtain statistical parameters characterizing the dispersion of the brightness distribution; calculating the local normal change rate of each pixel in the bending region using the three-dimensional geometric curvature parameters, and using the normal change rate as a constraint operator to perform spatial weight correction on the statistical variance; establishing a brightness response probability model about the viewpoint vector based on the corrected statistical variance, and generating an expected envelope surface describing the ideal brightness distribution of each pixel position under different observation orientations, as the photometric distribution model.

[0035] Using the pre-calibrated intrinsic and extrinsic parameter matrices of a multi-view camera array, the system projects the acquired multi-view image sequences onto a world coordinate system. Based on the geometric topological relationships obtained from 3D surface reconstruction, a discrete sampling grid with a sampling step size of 0.05 mm to 0.1 mm is established in space. For each spatial coordinate point in the grid, the system retrieves the corresponding pixel index from multiple original images using ray casting and extracts the grayscale value at non-integer coordinates using bilinear interpolation, thereby constructing a pixel brightness vector reflecting the change of that physical point with the observation angle.

[0036] Statistical calculations are performed on the pixel brightness vector at each coordinate point. By calculating the arithmetic mean of all elements in the vector and extracting the deviation of the grayscale value from each viewpoint relative to this mean, the initial pixel brightness variance at that coordinate point is obtained. This variance serves as a basic statistical parameter, preliminarily quantifying the brightness fluctuation characteristics of this spatial location under different observation orientations.

[0037] To compensate for the lighting and shadow interference caused by geometric deformation, the system introduces a spatial correction mechanism based on normal constraints. First, the system calculates the local normal change rate of each pixel in the bending region by performing a first-order differential operation on the normal vector within the local neighborhood. Specifically, a fixed-size neighborhood window (e.g., 5×5 pixels or 0.5mm in diameter) is taken on the 3D model around the current pixel. The average angle (angular displacement) between the normal vector of the center point within this window and the normal vectors of the surrounding points is calculated to characterize the degree of directional change at that point. For the bending ridge, due to the abrupt change in the geometric surface, the average angle increases significantly, thus quantifying the severity of the deformation. Subsequently, the system pre-constructs a mapping function that takes the normal change rate of each point as input and uses the average global normal change rate as the sensitivity threshold.

[0038] When the rate of change of the normal at a point exceeds a threshold, the mapping function outputs a gain value according to exponential growth logic based on the natural logarithm base, and uses a linear stretching operator to map this gain to a defined interval of 1.2 to 3.0, thereby generating a constraint operator. Within this interval, 1.2 corresponds to the transition region where the normal begins to deflect, while 3.0 corresponds to the ridge vertex region with the largest bending radius. This mapping function uses the ratio of the rate of change of the normal at the current point to a preset average threshold as its independent variable. When the ratio is greater than 1, the output gain exhibits a non-linear increasing trend with the increase of the ratio. The specific logic is as follows: when the rate of change is close to the threshold, the gain increases relatively slowly, remaining around 1.2; as the rate of change further increases (i.e., entering the deepest bend), the gain increases faster until it reaches the preset maximum upper limit of 3.0. This design is to reserve a larger tolerance space for brightness fluctuations at the bend apex, so that it is not misjudged as a defect. This value range is pre-calibrated based on the diffuse reflectance of the tab adapter material, ensuring that the weight distribution matches the physical reflection characteristics of the metal surface. The calibration process uses standard qualified parts (defect-free samples). First, the basic brightness variance of the flat area under standard illumination is measured; then, the natural halo or brushed texture generated at the maximum bend of the sample is manually observed, and the original variance multiple at that point is recorded. If the variance at the maximum halo is 2.5 times the basic variance, the upper limit of the gain of the mapping function is set to a value slightly higher than that (such as 3.0) to ensure that all normal process lighting and shadows can be included within the envelope of the photometric distribution model.

[0039] The generated constraint operator is applied as a multiplication operator to the initially calculated pixel brightness variance. In bending regions with drastic normal changes, the system artificially amplifies the allowable fluctuation range of variance in these regions due to the higher weighting coefficients output by the constraint operator. This correction effectively compensates for the natural brightness fluctuations caused by metal stretching and curvature abrupt changes, ensuring that phenomena such as "wire drawing" or "halo" generated by normal processes are included within a reasonable statistical envelope, thus avoiding false alarms caused by geometric morphology.

[0040] Finally, the system uses the observation azimuth unit vector of each camera as the input variable, and combines it with the corrected statistical variance to establish a brightness response probability model using a parameterized modeling method based on the second-order spherical harmonic function. Before fitting, the system first converts the three-dimensional azimuth vector of each camera relative to the sampling point into the horizontal azimuth angle and the vertical pitch angle. The second-order spherical harmonic model consists of 9 basic components, representing the light intensity distribution weights in different directions. The system uses the aforementioned corrected variance as a weighting factor. By minimizing the cumulative error between the measured brightness and the model's predicted brightness, it calculates the specific coefficient values ​​of these nine components, thereby obtaining a smooth brightness prediction surface that can cover all observation angles. During the fitting process, the system introduces an anti-outlier algorithm (such as the RANSAC algorithm) to remove pixel saturation extrema caused by specular reflection in the brightness vector. The RANSAC algorithm sets a consistency threshold during the iteration process, that is: if the measured gray value at a certain viewpoint deviates from the predicted value of the current fitted model by more than 20% of the total dynamic range of the image (for example, a deviation of more than 50 gray levels in an 8-bit image), then the point is determined to be an outlier point affected by strong specular light. The system will discard these points and use only the diffuse reflection data that conforms to the majority of viewing angles to lock in the final photometric model. By performing weighted least squares on the remaining valid sampling points, a smooth expected envelope surface describing the ideal brightness distribution of each pixel position under different viewing orientations is fitted. This envelope surface completely defines the photometric distribution characteristics of each pixel point, thus constituting the final photometric distribution model, providing a physical benchmark for subsequent accurate separation of specular reflection components and diffuse structure features.

[0041] By extracting pixel brightness vectors from multiple perspectives and combining them with the three-dimensional normal change rate for spatial weight correction, this method can objectively quantify the brightness fluctuation characteristics of metal surfaces under different observation orientations. By using the desired envelope surface fitted by the second-order spherical harmonic function, the specular reflection component and diffuse structural features are effectively separated, thereby reducing the interference of natural halos or textures caused by geometric deformation in the bending area on the detection results. This feature processing method based on physical priors alleviates the risk of false alarms caused by uneven ambient light field and part shape distortion, and enhances the feature consistency of the input deep learning network.

[0042] Furthermore, the process of separating the specular reflection component field and the diffuse structure feature field from the multi-view image sequence includes: calculating the photometric response residual of the measured pixel brightness in the multi-view image sequence relative to the expected envelope surface in the photometric distribution model; constructing an anisotropic threshold function in combination with the three-dimensional geometric curvature parameters, performing logical judgment on the photometric response residual, and extracting outlier pixels exceeding the threshold as the specular reflection component field reflecting the instantaneous strong light on the metal surface; stripping the specular reflection component field from the original multi-view image sequence, and performing cross-view consistency feature enhancement on the remaining pixel components, outputting a diffuse structure feature field characterizing the micro-texture and diffuse reflection structure of the metal surface.

[0043] First, pixel-by-pixel calculation of the photometric response residual is performed. Using a pre-calibrated camera intrinsic and extrinsic parameter matrix, the pixel coordinates of each frame of multi-view image are spatially aligned with the reconstructed 3D geometric surface. During the mapping process, the system calls the depth buffer algorithm to compare the distance from the 3D sampling point to the camera in real time along the same optical axis. Through occlusion culling logic, the measured grayscale value is extracted only for visible points that are within the camera's field of view and are not occluded by the shape of the part itself. Then, based on the current camera's observation azimuth vector, the ideal brightness expectation value corresponding to the point is retrieved from the pre-established spherical harmonic photometric distribution model. By calculating the absolute difference between the measured grayscale value and the expectation value, the initial photometric response residual of the pixel is generated.

[0044] To achieve accurate determination, the statistical variance of background noise is first extracted from the flat substrate area of ​​the part, and a basic residual tolerance value is set based on this (e.g., set to 3 times the background standard deviation). Subsequently, the system calculates the product of the Gaussian curvature and the mean curvature at the sampling point, and uses this as a deformation intensity index to construct a dynamic threshold function. In this embodiment, the system divides the surface into three levels for tiered management: in the flat area, the basic residual tolerance value is directly applied; in the transition area (when the deformation index begins to rise), the threshold is set to 1.5 times the basic value; in the ridge core area (when the deformation index reaches the preset maximum process deformation threshold), the threshold is increased to 2.5 to 3.0 times the basic value. This tiering mechanism ensures that at the ridge where the geometric change is most severe, the system can accommodate normal process halo by relaxing the tolerance.

[0045] In the threshold adjustment execution logic, an asymmetric hysteresis judgment mechanism is used to handle region switching. When the sampling point moves from the core area of ​​the ridge to the flat substrate on both sides and the curvature index gradually decreases, the threshold does not immediately decrease with the index. The system executes a "hysteresis protection" logic: only when the local curvature index drops below 80% of the trigger value corresponding to the current level, the system will lower the residual tolerance value back to the level of the previous level. This "fast acceleration and slow deceleration" switching logic effectively compensates for the morphology detection jitter caused by the small irregularities in metal processing and prevents false defect noise from being generated at the bending edge due to frequent threshold jumps.

[0046] After determining the threshold basic strength, the system further uses the feature vector (i.e., principal curvature direction) of the three-dimensional surface structure tensor to allocate spatial weights. The system extracts the direction with the smallest curvature change (along the ridge tangent direction) and the direction with the largest change (perpendicular to the ridge normal direction) at that point, and constructs an elliptical decision domain around the pixel. For the residual component along the ridge tangent direction, the system applies the aforementioned upgraded upper limit threshold for loose judgment to tolerate the natural brushed texture distributed along the bend. For the component in the normal direction perpendicular to the ridge, a strict basic threshold is restored in this direction through coordinate rotation transformation. This differentiated allocation of "loose in the longitudinal direction and strict in the transverse direction" enables the algorithm to keenly capture the tiny cracks that cut perpendicularly to the flow direction of metal grains, while filtering out the interference of normal physical properties along the bend direction.

[0047] The measured residuals are compared with the anisotropic threshold field constructed above. For any pixel that exceeds the threshold, the absolute value of the excess part is recorded as a residual intensity map, which constitutes the specular reflection component field. Subsequently, the system subtracts the intensity value pixel by pixel from the original multi-view image to obtain a preliminary image containing only diffuse reflection information. For the remaining pixels, the system uses cross-view geometric constraints to perform feature fusion. The fusion weight is determined according to the angle between the observation azimuth vector and the local surface normal: a cosine monotonically increasing function is used to give higher weights to the viewpoint with a larger angle (i.e., the oblique viewpoint), because these angles are least affected by specular reflection residue interference in physical properties. Through this weighted fusion based on physical angles, the system finally outputs a diffuse structure feature field with coherent texture, pure background and characterization of the micro-morphology of the metal surface.

[0048] Specifically, since the standard cosine function value exhibits a monotonically decreasing characteristic in the range of 0 to 90 degrees, in order to assign higher weight to the oblique viewing angle with a larger included angle, this embodiment constructs a monotonically increasing cosine function with respect to the included angle θ by inverting the cosine value. Its mathematical expression is as follows: W(θ) = k·(1-cosθ))+b Where θ is the angle between the azimuth vector of the current observation field and the local normal of the surface of the tab adapter, W(θ) is the feature fusion weight corresponding to this viewpoint; b is the basic weight bias coefficient, used to ensure that the fusion weight of the normal viewpoint (i.e., when θ is close to 0 degrees) is not zero, so as to retain the basic diffuse texture information that is not overexposed under this viewpoint. The value of b is pre-calibrated according to the basic diffuse reflectance of the tab adapter material, and the preferred value range is 0.1 to 0.3; k is the weight gain coefficient, used to control the magnification of the oblique viewpoint weight. Due to the high reflectivity of the tab adapter, the oblique viewpoint is least affected by specular reflection interference. Therefore, it is necessary to significantly improve its feature contribution by using the k value. The value of k is adaptively and dynamically adjusted according to the illumination intensity of the system ring light source and the surface roughness of the part, and the preferred value range is 0.8 to 1.5. In actual engineering deployment, technicians can conduct multi-view lighting tests on defect-free standard tab adapter samples, compare the saturation pixel ratio of images from each viewpoint, and then lock the optimal combination of k and b parameters.

[0049] By introducing an anisotropic threshold adjustment mechanism based on three-dimensional curvature compensation, this method effectively decouples specular reflection and diffusion features in complex bending regions. The stepped enhancement and hysteresis judgment logic reduces interference from natural halos caused by abrupt changes in geometric shape, lowering the risk of false alarms. Combined with weight allocation along the principal curvature direction, it suppresses ambient light field noise while preserving the microscopic grain texture of the metal surface to the maximum extent, thus providing consistent feature input for subsequent deep learning networks to identify minute physical damage.

[0050] Further, the process of generating the corrected texture map includes: determining the differential manifold structure of each sampling point in the diffuse structure feature field using the three-dimensional geometric curvature parameters, and calculating the local metric tensor characterizing the degree of surface stretching or compression; using the geometric center line of the bending region as a reference, calculating the projection transformation coefficients of each pixel point on the tangent plane using the local metric tensor, and thereby constructing the local tangent space mapping matrix of each pixel position; using the local tangent space mapping matrix to establish the nonlinear mapping relationship between each sampling point in the diffuse structure feature field in the three-dimensional surface coordinates and the two-dimensional Euclidean plane coordinates, and projecting the surface texture onto the standard two-dimensional coordinate system; establishing an orthogonal pixel grid in the two-dimensional Euclidean plane, resampling the projected feature components using a bilinear interpolation algorithm, filling the image pixel holes caused by the abrupt curvature change, and outputting the isotropic corrected texture map.

[0051] First, using the acquired Gaussian curvature, mean curvature, and surface normal direction, each sampling point in the diffuse structure feature field is locally parameterized. Specifically, a spherical neighborhood with a radius of 0.5 mm is taken at each three-dimensional sampling point. The discrete points in this neighborhood are fitted to a two-variable quadratic continuous surface equation using the least squares method. By taking the partial derivative of this equation, the first fundamental form coefficients (i.e., metric coefficients E, F, G) of the point in the two orthogonal tangent vector directions are calculated. These coefficients together constitute the local metric tensor, which serves as a physical index for measuring the stretching or compression ratio of the surface relative to the standard plane in three-dimensional space. For example, if the coefficient E is greater than 1, it is determined that there is physical stretching in the corresponding direction.

[0052] Using the curved ridge line (geometric centerline) of the bending region as the reference origin for spatial positioning, for each spatial sampling point to be corrected, the system employs a fast traversal method to calculate the shortest path from that point to the centerline along the 3D surface, i.e., the geodesic distance. Simultaneously, the system extracts the tangent vector of this path at the sampling point and performs Schmitt orthogonalization with the normal vector of that point, thereby constructing a 3×3 local tangent space mapping matrix specific to that point. This matrix rotates and aligns the 3D local coordinate system of that point, ensuring that two of its coordinate axes are parallel to the direction of the centerline and perpendicular to the cross-sectional direction of the centerline, respectively.

[0053] Using the generated mapping matrix, a nonlinear correspondence between surface coordinates and two-dimensional plane coordinates is established. In engineering implementation, a step-integration method along the geodesic is adopted: starting from the center line, stepping to both sides according to a fixed small arc length (such as 0.01 mm), and using the local metric tensor at each step point to correct the arc length (compensate for stretching or compression), converting it into linear displacement coordinates on the two-dimensional plane. Through this cumulative calculation, all sampling points on the three-dimensional surface are "tiled" into the standard two-dimensional Euclidean Cartesian coordinate system. This process ensures that the texture at the bend does not have its geometric topology (such as crack direction and grain arrangement) distorted after unfolding.

[0054] Within the two-dimensional standard coordinate system formed by projection, the system pre-establishes a set of orthogonally arranged pixel grids based on the physical resolution of the original camera (e.g., set to 0.02 mm / pixel). Since the projected points in areas that were originally severely stretched (such as the outer side of a bend) become sparse after the surface is unfolded, resulting in pixel holes, a reverse mapping search logic is employed: traversing each pixel target point in the two-dimensional grid, and using a bilinear interpolation algorithm, searching for the four nearest original sampling points in the non-uniform feature field after projection. By weighted fusion of the feature values ​​(grayscale or polarization components) of these four points, the system smoothly fills all holes and compensates for image brightness fluctuations caused by uneven sampling rates. The final output corrected texture map has a globally uniform scale (isotropic), providing standard image input for subsequent deep learning networks.

[0055] The specific implementation process of the weighted fusion includes: First, the system selects a pixel to be generated (i.e., the target point) in a preset two-dimensional orthogonal grid. Using the established local tangent space mapping matrix, the system determines the precise position of the target point in the projected feature field through inverse mapping logic. Since the point distribution after surface unfolding has nonlinear characteristics, this position is usually represented as a set of floating-point coordinates. Subsequently, the system searches for and locks the four closest original sampling points around these floating-point coordinates. These four points form a minimum rectangular unit surrounding the target point in the topology. The weight calculation logic is based on the relative spatial displacement of the target point within this rectangular unit. The system extracts the horizontal and vertical offsets of the target point relative to the upper left corner vertex of the rectangle and divides these two displacement values ​​by the original sampling interval to obtain two normalized scaling factors between 0 and 1. These two scaling factors constitute the core source of the weights: the closer the target point is to a certain sampling point in physical distance, the higher the weight ratio of that sampling point in the subsequent fusion.

[0056] The specific fusion operation adopts a two-level linear superposition path. In the first-level fusion, the system uses a horizontal scaling factor to perform feature weighting on a pair of neighboring points above and below the rectangular unit, respectively, to obtain two intermediate feature values ​​on the horizontal axis of the target point. In the second-level fusion, the system uses a vertical scaling factor to perform weighted superposition on these two intermediate feature values ​​again. Through this horizontal to vertical stepwise fusion, the feature information of the four original sampling points (such as polarization degree or grayscale residual) is completely transformed into the final value of the target pixel.

[0057] This weighted fusion mechanism ensures the continuity of the generated corrected texture map in terms of physical properties. By utilizing the feature contributions of the four surrounding points, the system can smoothly fill the projection holes caused by drastic changes in bending rate and automatically compensate for image fluctuations caused by inconsistent sampling rates. The final output isotropic corrected texture map contains precisely calculated physical features for each pixel, thus providing a standardized data source for subsequent deep defect recognition networks to extract features at a unified morphological scale.

[0058] By constructing a tangent space mapping matrix using local metric tensors, the diffuse structural feature field on a three-dimensional curved surface is nonlinearly projected onto a standard two-dimensional plane coordinate system, effectively compensating for radial and axial distortions in the image caused by part bending. This process, combined with bilinear interpolation weighted fusion, fills pixel holes caused by abrupt curvature changes, ensuring the topological consistency of the metal texture and potential defects before and after projection.

[0059] Further, the process of determining candidate defect pixels based on the singular values ​​of the gradient field includes: performing multi-scale Gaussian convolution on the corrected texture map, calculating the second-order partial derivatives of each pixel coordinate point in the horizontal, vertical, and diagonal directions, and constructing a Hessian matrix representing the local gray-level change rate; constructing a local structure tensor using the elements of the Hessian matrix, and performing spatial smoothing integration on the structure tensor using a preset window to fuse the metal grain flow direction features of adjacent regions; performing singular value decomposition on the integrated structure tensor, extracting the principal singular values ​​reflecting the texture energy intensity and the feature vectors reflecting the consistency of the texture direction, and constructing a gradient field; calculating the anisotropy confidence of each pixel position in the gradient field, and determining the pixel positions with confidence scores lower than a preset threshold and whose principal singular values ​​undergo abrupt changes as candidate defect pixels, thereby identifying local morphological anomalies that disrupt the continuity of the metal grain flow direction.

[0060] First, multi-scale Gaussian convolution processing is performed on the generated corrected texture map. In the specific implementation, a set of Gaussian kernels with standard deviations of 1.0, 2.0 and 4.0 are preset for noise suppression. Then, using the discrete difference operator, the second derivatives of each pixel in the horizontal and vertical directions, as well as the mixed partial derivatives at the intersection of the horizontal and vertical directions, are calculated at each scale. By arranging the horizontal second derivatives, vertical second derivatives and mixed partial derivatives according to the correspondence in the standard space, a 2×2 Hessian matrix is ​​constructed for each pixel position to capture the curvature change rate of local grayscale at the micro level.

[0061] Next, a local structure tensor is constructed using the elements in the Hessian matrix. In the specific operation, the second-order response components in the Hessian matrix are squared and cross-multiplied to generate the initial tensor components. In order to capture the macroscopic "grain flow direction", the system introduces a 7×7 or 9×9 square integration window and performs spatial smoothing integration (i.e., weighted average) on the tensor components of all pixels in the window. This step effectively filters out single-point photosensitive artifacts on the metal surface by fusing the geometric responses of adjacent pixels, enabling the algorithm to extract the distribution pattern that represents the overall trend of grain arrangement.

[0062] Singular value decomposition is performed on the integrated local structure tensor. For the symmetric matrix of each pixel, the system calculates two non-negative singular values ​​and their corresponding eigenvectors. The larger principal singular value reflects the energy intensity of the local texture at that point (i.e., the clarity of the brushed metal texture), while the corresponding eigenvector indicates the principal direction of the texture energy distribution (i.e., the physical extension direction of the metal grains). Using these eigenvectors, the system constructs a continuous directional gradient field in the image space, which fully describes the texture topology of the corrected metal surface.

[0063] Finally, the defect candidate point determination is performed, and the anisotropy confidence of each position in the gradient field is calculated. This confidence is achieved by analyzing the degree of difference between the primary and secondary singular values: the ratio of the difference between the two singular values ​​to the sum of the two singular values ​​is calculated. When the ratio is close to 1, it indicates that the point has extremely strong directional consistency and belongs to the normal grain flow direction. When the ratio is lower than a preset threshold (0.4 in this embodiment), it indicates that the directionality at that point is disordered. Pixel positions with such low confidence and a significant decrease or abrupt change in the primary singular value relative to the mean of the surrounding neighborhood are identified as defect candidate pixels. This logic can accurately identify the interruption of grain flow direction caused by microcracks and indentations, thereby locking in potential geometric damage areas.

[0064] By constructing a local structure tensor and extracting the directional features of the gradient field using singular value decomposition, the coherence of grain flow direction can be quantified from a complex background. This determination method based on anisotropic confidence effectively distinguishes normal metal wire drawing texture from micro-cracks or indentations that disrupt the flow direction coherence, and reduces artifact interference caused by ambient light and shadow.

[0065] Furthermore, the multi-view spatial consistency verification includes: mapping the Gaussian curvature and average curvature in the three-dimensional geometric curvature parameters to physical stress concentration probabilities, generating the stress distribution prior mask, which is used to assign higher confidence weights to candidate pixels in high-curvature bending regions; using the relative pose matrix of the camera device when acquiring multi-view image sequences, calculating the epipolar equation between each observation viewpoint, and using it as the epipolar geometric space constraint criterion; performing a one-dimensional search along the epipolar line in images of adjacent viewpoints, calculating the cross-view feature matching degree of the defect candidate pixels, eliminating pseudo-defect points that cannot satisfy geometric projection constraints in multiple viewpoints or whose comprehensive score after combining anisotropic confidence is lower than a set threshold, and retaining the real defect pixel region with multi-view spatial consistency.

[0066] First, the Gaussian curvature and mean curvature at each sampling point on the three-dimensional surface are extracted. Since the stress concentration depends mainly on the absolute value of surface deformation rather than the direction of bending, the two curvature parameters are first processed to be absolute values. A comprehensive curvature index is generated by weighted summation. In this embodiment, the weight allocation logic is based on the sensitivity of the metal material to different deformation modes: Gaussian curvature mainly reflects the overall stretching or contraction of the surface (inherent geometric characteristics), while mean curvature mainly reflects the degree of bending of the surface (external geometric characteristics). For the bending process of the tab adapter, the system sets the weight of mean curvature to be higher than that of Gaussian curvature (in this embodiment, the weight of mean curvature is set to 0.7 and the weight of Gaussian curvature is 0.3), because the cracks at the bending point are mainly induced by normal bending stress. By multiplying the absolute values ​​of the two curvatures with their respective weights and summing them, the system obtains a comprehensive curvature index that can quantify the local total geometric deformation energy.

[0067] First, technicians need to perform deformation pattern recognition on the processing technology of the parts to be inspected. At this stage, it is necessary to distinguish whether the surface of the part is in "external bending" or "internal distortion". The average curvature describes the degree of folding of the surface in three-dimensional space, corresponding to normal bending strain; while Gaussian curvature describes the internal stretching or contraction of the surface, corresponding to the thinning or tearing of the material. For parts such as tab adapters, where bending is the core process, since cracks are mainly induced by normal bending stress, the average curvature needs to be given a higher weight in the process (such as 0.7 in this embodiment) to ensure the algorithm's sensitivity to deformation at the bending ridge.

[0068] Specific weight values ​​are determined through simulation correlation experiments. The stress distribution cloud map of the part under standard machining conditions is obtained using finite element analysis. Simultaneously, the average curvature field and Gaussian curvature field of the same model are calculated using algorithms. By calculating the Pearson correlation coefficients between these two geometric parameter fields and the actual physical stress field, the weight ratio can be scientifically derived. For example, when the simulation shows that the correlation between bending stress and stress concentration is significantly higher than that of material tension, the weight coefficient H should be increased proportionally. Subsequently, background noise calibration and stress trigger threshold setting are performed. Curvature parameters are extracted from the flat substrate area of ​​the part as environmental background noise, and a stress trigger threshold is set. Only when the comprehensive curvature index exceeds this threshold is the area considered to have a stress concentration risk. For materials with poor ductility and prone to cracking, the trigger threshold should be lowered and the weight sensitivity appropriately increased.

[0069] Finally, a nonlinear mapping from geometric exponent to physical probability is completed. The calculated comprehensive curvature exponent is mapped to the probability value in the [0, 1] interval through linear stretching or step enhancement logic, thereby generating a stress distribution prior mask. This mask will serve as the spatial attention weight of the subsequent deep learning network, guiding the network to strengthen the response to high curvature and high stress regions in the feature extraction layer, and realizing the accurate guidance of physical prior to algorithm reasoning.

[0070] Subsequently, a nonlinear mapping is performed to convert the index into a physical stress distribution probability. A stress trigger threshold is preset, which corresponds to the normal geometric fluctuations in the flat substrate area of ​​the part. When the comprehensive curvature index at a certain point is lower than the threshold, the mapped stress probability is set to an extremely low value or zero. When the index exceeds the threshold, linear stretching or stepped enhancement logic is used to map the index to a probability range between 0 and 1. During the mapping process, the larger the comprehensive curvature index, the higher the corresponding stress distribution probability, thus forming a probability peak close to 1 at the bending ridge position. These probability values ​​are rendered as a stress distribution prior mask aligned with the original image space. This mask records the stress concentration degree at each physical location in the form of pixel grayscale values. The larger the value, the higher the probability of physical damage in that area.

[0071] Using the relative pose matrices and camera intrinsic parameters of each viewpoint obtained during the initialization phase, the system calculates the fundamental matrix and derives the corresponding epipolar equation for any two adjacent observation viewpoints. When a candidate pixel for a defect is locked in the main view, the system uses the depth search range obtained from the 3D reconstruction of that point (determined by the maximum and minimum elevation of the part surface) to determine a finite-length epipolar line segment on the adjacent view plane, instead of a straight line across the entire image. This line segment represents all possible legal projection positions of the physical point from another viewpoint, compressing the feature search into a precise one-dimensional line segment search, which constitutes the geometric benchmark for spatial consistency verification.

[0072] For candidate defect pixels in the main view, the system searches along epipolar lines generated in adjacent views. Considering the difference in reflectivity of the metal surface under different viewing angles, the system extracts a 5×5 pixel window around the candidate point and uses a mean-free normalized cross-correlation algorithm to match it with the candidate region on the epipolar line. This algorithm eliminates the influence of uneven illumination between viewing angles by subtracting the window mean and dividing by the standard deviation. The system records the coordinates of the highest matching score and calculates the vertical deviation distance from that point to the theoretical epipolar line. If the deviation is less than 1.5 pixels and the matching score exceeds a preset ratio, the cross-view feature matching degree is judged to be high.

[0073] Finally, the system executes a comprehensive scoring logic. The system uses the cross-view feature matching degree as the base score, adds a deduction item for geometric projection error, and finally multiplies it by the confidence weight provided by the stress distribution prior mask. For outliers that appear only in a single view (such as instantaneous specular reflection) or pixels that cannot be consistently matched on the epipolar line, the system judges them as pseudo-defect points and removes them. Only regions that meet the epipolar geometric constraints in at least three views and have a comprehensive score of more than 75 points (out of 100) are defined as real defect pixel regions with consistent multi-view spatiality. These verified regions will be used as the final reliable feature input to enter the deep learning network for classification and recognition.

[0074] By transforming geometric curvature into a priori physical stress distribution, quantitative weighting of high-risk areas such as folded ridges is achieved. Combined with epipolar geometric constraints, this method can effectively eliminate false defects caused by instantaneous reflections or dust, ensuring that the detection features have cross-view consistency. This significantly enhances the physical credibility of the input neural network data, making the model more focused on the identification of real geometric damage.

[0075] Further, the output surface defect detection result includes: using the real defect pixel region as a spatial guiding channel, and splicing and fusing it with the corrected texture map containing local topological information in the channel dimension to construct a multi-channel feature input tensor; inputting the multi-channel feature input tensor into the pre-trained deep defect recognition network, and injecting the stress distribution prior mask as a spatial attention weight matrix into the feature extraction layer of the deep defect recognition network to guide the network to strengthen the feature response to high stress concentration areas. Specifically, the spatial dimension of the stress distribution prior mask is scaled to be consistent with the size of the feature map output by the feature extraction layer through a bilinear interpolation algorithm to guide the network to improve the perception intensity of local morphological anomalies; using the multi-scale feature pyramid in the deep defect recognition network to extract defect semantic features under different receptive fields, and after passing through a classification and regression head network, outputting specific category labels and localization bounding boxes representing microcracks, indentations, or scratches on the surface of automotive parts.

[0076] First, the system performs dimensional alignment and unit unification of the input features. It extracts the isotropic calibrated texture map generated in the previous steps as the main visual channel and linearly normalizes its pixel grayscale values ​​to the range of 0 to 1. Simultaneously, it uses the real defect pixel regions determined by multi-view verification (binarized mask image, where potential defect pixels are 1 and background pixels are 0) as the spatial guide channel. The system concatenates these two channels along the channel dimension to construct an input tensor of dimension H×W×2. This construction method ensures that the first convolutional kernel of the convolutional neural network, when performing sliding window calculations, can simultaneously perceive the surface micro-texture and geometrically verified candidate damage locations through cross-channel correlation.

[0077] The network used in this embodiment is a physical perception convolutional neural network based on an improvement of ResNet-50. This network uses four sets of standard residual blocks as the backbone feature extractor. Its core improvement is that after the output of the first feature extraction stage of the backbone network (i.e., the first residual block group Stage1), a dedicated spatial attention injection module is deployed. This module has an independent bypass input interface for receiving a stress distribution prior mask that has been scaled by bilinear interpolation to be completely consistent with the spatial size of the Stage1 feature map. In addition, the network ends by constructing a multi-scale feature pyramid through 1x1 convolutional layers and upsampling operations to achieve the layer-by-layer fusion of high-level strong semantic information and low-level high-resolution detailed features.

[0078] Specifically, the input tensor space size of the deep defect recognition network is consistent with the corrected texture map, denoted as H×W, and the number of channels is 2. The first channel is the corrected texture map after linear normalization, and the second channel is the binary mask map of the real defect pixel area. For general industrial vision applications, H and W are preferably set to even values ​​in the range of 512 to 2048 pixels.

[0079] The backbone network adopts a 50-layer residual network structure, including four groups of residual blocks: conv1, Stage1 to Stage4. The number of channels in the output feature maps of Stage1 to Stage4 are 256, 512, 1024 and 2048, respectively. The spatial attention injection module takes the feature map output by Stage1 as the main branch input and a stress distribution prior mask, which is scaled to the same spatial size as the feature map of Stage1 by bilinear interpolation, as the bypass input. The stress distribution prior mask is copied and expanded to 256 channels in the channel dimension and is multiplied element-wise with the feature map of Stage1 to obtain the first-stage feature map after adding spatial attention weights.

[0080] The multi-scale feature pyramid takes the output feature maps of Stage2, Stage3, and Stage4 as input, reduces the dimensionality to a uniform 256 channels through 1×1 convolution, and constructs pyramid feature maps of three scales by top-down progressive upsampling and element-wise addition. The classification head and regression head are composed of two layers of 3×3 convolution and one layer of 1×1 convolution, respectively. The pyramid feature map of each scale is used as input to output the corresponding scale's category prediction and location regression results.

[0081] During the forward propagation of the network, the spatial attention injection module executes specific feature enhancement logic. Since the feature map output by Stage1 has multi-channel attributes (such as 256 channels), while the input stress mask is single-channel, the single-channel mask is copied and expanded in the channel dimension using the tensor broadcast mechanism to make its dimension completely aligned with the feature map. To prevent the features of non-stress areas from being directly masked, which would cause the model to fail to converge, the system executes weight mapping logic: the normalized stress probability value M is converted into a weight coefficient W through linear offset. The specific logic is 1.0 + sensitivity coefficient × M (in this embodiment, the sensitivity coefficient is set to 1.2). Subsequently, the system performs a pixel-wise dot product with the feature map. Through this operation, the system forcibly guides the network to artificially amplify the activation intensity of neurons in physically high-risk areas such as bent ridges, thereby significantly improving the ability to capture microscopic abnormal flow directions in complex background textures.

[0082] The performance of this deep recognition network is ensured through the following training process: Addressing the scarcity of microcrack samples in industrial settings, the system utilizes a generative adversarial network (GAN) to extract typical defect textures. During training set construction, the generated defect textures are controlled and implanted into the high-stress mask regions of normal samples. A Poisson fusion algorithm is then used to eliminate splicing marks, ensuring that the training samples conform to the physical failure logic of metallic materials. Specifically, the defect texture size output by the GAN is preferably set to a range of 16×16 to 64×64 pixels, with a grayscale range consistent with the grayscale range of the corrected texture map. When constructing training samples, for each defect-free normal sample image, 1 to 5 center locations are randomly selected within areas where the prior mask value of the stress distribution is greater than 0.7. Defect texture blocks with an area not exceeding 5% of the entire image area are implanted, and the edges are smoothly spliced ​​within a transition band of 3 to 5 pixels wide using a Poisson fusion algorithm. These parameter ranges ensure that the number and size of the synthesized defects match the actual working conditions, preventing deviations in the training sample distribution due to over-implantation. A usable training set can be obtained by adjusting the number and size of the defect blocks within this range.

[0083] The classification head uses Focal Loss to address the imbalance between minor defects and a large number of background samples by adjusting the weight factors. The regression head uses the Smooth L1 loss function to improve the robustness of localizing bounding boxes. The training hyperparameters employ stochastic gradient descent with a driving term, and the initial learning rate is set to 1×10⁻⁶. -4 The batch size was set to 32. Data augmentation strategies such as rotation, brightness jitter, and simulated wire-drawing noise were added during training until the average accuracy on the validation set reached the preset delivery standard.

[0084] During inference, the trained network feeds the fused multi-scale feature map into the decoupled classification and regression heads. The classification head outputs the specific category label and confidence score for whether the current region belongs to a microcrack, indentation, or scratch; the regression head predicts the center coordinates, width, and height of the defect in the two-dimensional correction plane. The system first uses a non-maximum suppression algorithm to remove redundant boxes with an overlap rate higher than 0.45. Finally, the system uses the inverse function of the aforementioned local tangent space mapping matrix to remap the detection results in the two-dimensional coordinate system back to the three-dimensional geometric coordinate system of the part. Ultimately, the system outputs the specific category, physical location, and three-dimensional bounding box results representing the surface damage of the component, achieving high-reliability quality inspection with deep coupling of physical consistency and deep learning inference.

[0085] By fusing physical prior masks with calibrated texture maps through multiple channels, deep coupling between visual features and stress distribution is achieved. A spatial attention mechanism guides the network to focus on high-risk bending areas, enhancing the model's feature response to microscopic damage. Combining multi-scale feature pyramids and coordinate inverse calculation logic, this process effectively improves the reliability of defect localization in complex metallic texture backgrounds while maintaining physical consistency.

[0086] By deeply coupling 3D geometric topography and photometric physical models, the problem of feature decoupling in complex bending areas of high-reflectivity metal parts is effectively solved. Using differential manifold unrolling and metric tensor correction, standardized generation of isotropic texture maps is achieved, eliminating the interference of geometric distortion on defect feature extraction from the bottom layer. Combined with a spatial attention mechanism based on stress distribution priors and multi-view epipolar geometric constraints, ambient light and noise interference are accurately filtered in both physical and geometric dimensions. This significantly improves the sensitivity of deep learning networks in recognizing micro-cracks, indentations, and scratches, enhancing the robustness of complex topography detection and providing a reliable guarantee of physical consistency for high-precision industrial automated quality inspection.

[0087] Example 2:

[0088] At the inspection station, the system places the aluminum alloy adapter piece under test under a ring light source with four-way independent control. The camera device captures four original images at high speed from directly above the part, each excited by light from different directions. At this time, the algorithm backend retrieves the three-dimensional geometric guidance parameters of the part and uses a local tangent space mapping matrix to resample the image of the bent R-angle area. The physical effect of this step is to "flatten" the curved surface, which originally suffered from perspective shrinkage and shadow occlusion due to bending, into a standardized, isotropic, and uniformly sized texture map in digital space. This ensures that subsequent inspection is performed on a unified planar scale, eliminating visual misleading effects caused by geometric distortion.

[0089] The system automatically associates the original CAD design model of the adapter piece, extracts the differential geometric properties of the bending parts, and generates a stress distribution prior mask aligned with the texture map space by performing absolute value weighted calculation on the Gaussian curvature and average curvature of each sampling point on the surface. Visually, this mask marks the probability distribution of bending ridges and areas of severe stress, like a "heat map". This mask will then serve as a physical prior signal to tell the detection algorithm which areas are the most vulnerable and most prone to microcracks, thus establishing a "key prevention and control" mechanism in the full-map search.

[0090] The algorithm performs microscopic flow direction analysis on the corrected texture map. Since the aluminum alloy surface has a regular brushed grain texture, the system uses the Hessian matrix to capture the second-order rate of change of local gray levels and extracts the principal direction energy of the texture through singular value decomposition. When a sudden change in the principal singular value of a certain region is detected, and the confidence level of its texture direction consistency decreases significantly, the system determines that the metal grain flow direction at that location has been physically interrupted. The system then marks these local morphological anomalies that disrupt the continuity as candidate defect pixels, achieving the goal of quickly locating suspected damaged areas from massive background textures.

[0091] To eliminate specular reflection artifacts and random dust interference commonly found on highly reflective aluminum alloy surfaces, the system incorporates a secondary verification using a side camera. Utilizing a calibrated camera pose matrix, the system calculates the epipolar equation between the main and side views. For candidate points identified in the previous step, the system performs one-dimensional feature matching along the corresponding epipolar segment in the side view. If a candidate point cannot find a corresponding feature satisfying the geometric projection rules in the side view, the system determines it to be a false defect caused by instantaneous reflection and discards it. Only points that align in both views and are located precisely within the high-stress prior region are confirmed as genuine defect areas.

[0092] Finally, the system performs channel-wise concatenation of the corrected texture map and the verified defect area map, and inputs it into the pre-trained improved deep recognition network. At this point, the stress distribution prior mask generated in the second stage is injected into the intermediate feature layer of the network as a spatial attention weight matrix. Through pixel-wise multiplication in residual form, the network is guided to "automatically focus" on the micro-anomalies at the bending ridge. The network ultimately outputs the category label and location bounding box for the "micro-crack" based on the multi-scale feature pyramid. Before outputting the results, the system uses the inverse function of the mapping matrix to remap the center of the two-dimensional bounding box back to the three-dimensional physical coordinate system of the part, providing a precise physical gripping position for the downstream rejection mechanism, thus completing the entire automated quality inspection closed loop.

[0093] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for detecting surface defects of electrode adapter plates based on deep learning, characterized in that, include: Obtain multi-view image sequences of the bending area of ​​the electrode adapter piece and the corresponding three-dimensional geometric curvature parameters; Calculate the pixel brightness variance of the multi-view image sequence at the same spatial coordinates, establish a photometric distribution model, and separate the specular reflection component field and the diffuse structure feature field from the multi-view image sequence; A local tangent space mapping matrix is ​​constructed using the three-dimensional geometric curvature parameters, and the diffuse structural feature field is projected from the surface space to the two-dimensional Euclidean plane to generate a corrected texture map. The second-order partial derivative of the corrected texture map is calculated and an anisotropic structure tensor is constructed. The gradient field reflecting the flow direction of metal grains is extracted, and the candidate pixel points of defects are determined according to the singular values ​​of the gradient field. Based on the three-dimensional geometric curvature parameters, a stress distribution prior mask and epipolar geometric space constraint criteria are constructed, and the prior mask and the constraint criteria are used to perform multi-view spatial consistency verification on the defect candidate pixels. The verified pixel regions are fused with the corrected texture map and input into a pre-trained deep defect recognition network to output surface defect detection results.

2. The method for detecting surface defects of tab adapter plates based on deep learning according to claim 1, characterized in that, The multi-view image sequence was acquired by an equidistant ring array deployed above the bending area of ​​the tab adapter, including: The ridge-line frontal view, perpendicular to the central normal of the arc-shaped ridge in the bending region, is used to capture the strong reflected brightness field at the bend apex; the flank symmetrical view, an oblique observation angle symmetrically distributed on both sides of the arc-shaped ridge, is used to obtain the microscopic grain flow characteristics of the bend transition slope; the root grazing view, a low-elevation observation position located at the junction of the bending region and the flat substrate, highlights the microcracks and indentations caused by stress concentration through the shadow enhancement effect formed by grazing light.

3. The method for detecting surface defects of a tab adapter plate based on deep learning according to claim 2, characterized in that, The process of acquiring the multi-view image sequence and the corresponding three-dimensional geometric curvature parameters includes: simultaneously acquiring sub-images of the tab adapter in multiple orthogonal polarization states by configuring polarization modulation elements at the front end of the camera device at each observation position; each frame of the multi-view image is processed by multi-channel fusion to generate a composite physical feature map containing polarization degree and phase delay, which is used to distinguish the specular bright spot area from the actual metal surface micro-damage by utilizing the optical anisotropy of the metal surface under different bending curvatures; constructing a rough depth manifold of the bending region by utilizing the parallax relationship between the symmetrical viewpoints of the flanks, and inverting the surface normal vector by using the polarization component under the frontal viewpoint of the ridge to complete the depth manifold; calculating the Gaussian curvature and average curvature at each pixel coordinate point by performing differential geometric operations on the reconstructed three-dimensional surface, which are used as the three-dimensional geometric curvature parameters.

4. The method for detecting surface defects of a tab adapter plate based on deep learning according to claim 1, characterized in that, The process of establishing the photometric distribution model includes: extracting pixel brightness vectors at the same spatial coordinates in the multi-view image sequence, calculating the pixel brightness variance at each coordinate point to obtain statistical parameters characterizing the dispersion of the brightness distribution; calculating the local normal change rate of each pixel in the bending region using the three-dimensional geometric curvature parameters, and using the normal change rate as a constraint operator to correct the spatial weight of the variance; establishing a brightness response probability model about the viewpoint vector based on the corrected statistical variance, and generating an expected envelope surface describing the ideal brightness distribution of each pixel position under different observation orientations, as the photometric distribution model.

5. The method for detecting surface defects of a tab adapter plate based on deep learning according to claim 1, characterized in that, The process of separating the specular reflection component field and the diffuse structure feature field from the multi-view image sequence includes: calculating the photometric response residual of the measured pixel brightness in the multi-view image sequence relative to the expected envelope surface in the photometric distribution model; constructing an anisotropic threshold function in combination with the three-dimensional geometric curvature parameters, performing logical judgment on the photometric response residual, and extracting outlier pixels exceeding the threshold as the specular reflection component field reflecting the instantaneous strong light on the metal surface; stripping the specular reflection component field from the original multi-view image sequence, and performing cross-view consistency feature enhancement on the remaining pixel components, outputting a diffuse structure feature field characterizing the micro-texture and diffuse reflection structure of the metal surface.

6. The method for detecting surface defects of a tab adapter plate based on deep learning according to claim 1, characterized in that, The process of generating the corrected texture map includes: determining the differential manifold structure of each sampling point in the diffuse structure feature field using the three-dimensional geometric curvature parameters, and calculating the local metric tensor characterizing the degree of surface stretching or compression; using the geometric center line of the bending region as a reference, calculating the projection transformation coefficients of each pixel point on the tangent plane using the local metric tensor, and thereby constructing the local tangent space mapping matrix of each pixel position; using the local tangent space mapping matrix to establish the nonlinear mapping relationship between the three-dimensional surface coordinates and the two-dimensional Euclidean plane coordinates of each sampling point in the diffuse structure feature field, and projecting the surface texture onto the standard two-dimensional coordinate system; establishing an orthogonal pixel grid in the two-dimensional Euclidean plane, resampling the projected feature components using a bilinear interpolation algorithm, filling the image pixel holes caused by the abrupt curvature change, and outputting the isotropic corrected texture map.

7. The method for detecting surface defects of a tab adapter plate based on deep learning according to claim 1, characterized in that, The process of determining candidate defect pixels based on the singular values ​​of the gradient field includes: performing multi-scale Gaussian convolution on the corrected texture map, calculating the second-order partial derivatives of each pixel coordinate point in the horizontal, vertical, and diagonal directions, and constructing a Hessian matrix representing the local gray-level change rate; constructing a local structure tensor using the elements of the Hessian matrix, and performing spatial smoothing integration on the structure tensor using a preset window to fuse the metal grain flow direction features of adjacent regions; performing singular value decomposition on the integrated structure tensor, extracting the principal singular values ​​reflecting the texture energy intensity and the feature vectors reflecting the consistency of the texture direction, and constructing a gradient field; calculating the anisotropy confidence of each pixel position in the gradient field, and determining the pixel positions with confidence scores lower than a preset threshold and whose principal singular values ​​undergo abrupt changes as candidate defect pixels, thereby identifying local morphological anomalies that disrupt the continuity of the metal grain flow direction.

8. The method for detecting surface defects of a tab adapter plate based on deep learning according to claim 1, characterized in that, The multi-view spatial consistency verification includes: mapping the Gaussian curvature and mean curvature in the three-dimensional geometric curvature parameters to physical stress concentration probabilities, generating the stress distribution prior mask, using the relative pose matrix of the camera device when acquiring multi-view image sequences to calculate the epipolar equation between each observation viewpoint, and using it as the epipolar geometric space constraint criterion; performing a one-dimensional search along the epipolar line in the images of adjacent viewpoints, calculating the cross-view feature matching degree of the defect candidate pixels, eliminating pseudo-defect points that cannot satisfy the geometric projection constraints in multiple viewpoints or whose comprehensive score after combining anisotropic confidence is lower than a set threshold, and retaining the real defect pixel region with multi-view spatial consistency.

9. The method for detecting surface defects of a tab adapter plate based on deep learning according to claim 1, characterized in that, The output surface defect detection result includes: using the real defect pixel area as a spatial guiding channel, and splicing and fusing it with the corrected texture map containing local topological information in the channel dimension to construct a multi-channel feature input tensor; inputting the multi-channel feature input tensor into a pre-trained deep defect recognition network, the deep defect recognition network including a backbone network, a spatial attention injection module, a multi-scale feature pyramid, and a classification and regression head network; the backbone network includes a first feature extraction stage, a second feature extraction stage, a third feature extraction stage, and a fourth feature extraction stage connected in sequence; the first feature extraction stage extracts the multi-channel feature input tensor and outputs a first feature map; the spatial attention injection module scales the spatial size of the stress distribution prior mask to be consistent with the first feature map, and in the channel... After dimensional expansion, element-wise multiplication is performed with the first feature map to generate a weighted feature map with injected attention weights; the second feature extraction stage receives the weighted feature map and performs feature extraction, outputting a second feature map; the third feature extraction stage receives the second feature map and performs feature extraction, outputting a third feature map; the fourth feature extraction stage receives the third feature map and performs feature extraction, outputting a fourth feature map; the multi-scale feature pyramid receives the second, third, and fourth feature maps, and fuses them to construct a multi-scale pyramid feature map through convolutional dimensionality reduction and top-down stepwise upsampling and element-wise addition; the multi-scale pyramid feature map is input into the classification and regression head network, outputting specific category labels and localization bounding boxes representing microcracks, indentations, or scratches on the surface of the tab adapter.