Aluminum bar surface defect visual detection and grading classification method
By generating a three-dimensional height map of the aluminum rod surface using multi-frequency sinusoidal structured light and triaxial Doppler vibration measurement technology, and combining it with a dual-stream Transformer network and multimodal information fusion, the problems of imaging quality and feature extraction in aluminum rod surface defect detection are solved, achieving high-precision defect classification and scientific evaluation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- YIZHENG HAITIAN ALUMINIUM IND CO LTD
- Filing Date
- 2026-04-01
- Publication Date
- 2026-07-07
AI Technical Summary
Existing aluminum rod surface defect detection technologies suffer from limited imaging quality under high-speed rotation and strong specular reflection environments. Fine features are easily lost, and the single feature extraction dimension leads to a high confusion rate of similar defects. The evaluation mechanism lacks physical basis and cannot achieve high-precision three-dimensional reconstruction and scientific classification.
Image sequences are acquired using multi-frequency sinusoidal structured light, and subpixel-level motion compensation is performed using a triaxial Doppler vibration measurement device. A three-dimensional height map is generated by multi-frequency difference phase unfolding. A dual-stream Transformer network is used to fuse the features of the two-dimensional grayscale image and the three-dimensional height map. Multimodal information fusion is performed by combining eddy current and ultrasonic detection, and the severity of defects is calculated based on the three-dimensional height map.
It achieves high-resolution 3D defect reconstruction in complex optical environments, significantly improving the accuracy and classification confidence of defect detection, providing scientifically based grading results, and offering reliable decision support for downstream production processes.
Smart Images

Figure CN122346804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial visual inspection technology, specifically a method for visual inspection and classification of surface defects in aluminum rods. Background Technology
[0002] Aluminum rods are widely used in high-end manufacturing fields such as building profiles, automotive parts, and aerospace structural components. Their surface quality directly affects the formability, corrosion resistance, and mechanical properties of downstream processes. However, current online visual inspection of aluminum rod surfaces faces three major technical challenges: First, limited imaging quality and easy loss of fine features. Under the dual constraints of high-speed rotation and strong specular reflection of the aluminum rod, existing fixed-parameter light source-camera systems cannot achieve consistent high-contrast imaging across the entire rod surface. Environmental interference and motion blur easily mask minute surface defects.
[0003] Second, the feature extraction dimension is limited, leading to a high confusion rate with similar defects. Defects such as pores and indentations, inclusions and oxidation spalling have highly similar morphologies on a two-dimensional plane. Existing conventional convolutional neural networks have limited ability to extract features from these similar defects, resulting in a persistently high false positive rate and seriously affecting the reliability of detection.
[0004] Third, the evaluation mechanism lacks a physical basis and cannot scientifically guide production. Existing defect evaluation methods often only reach the level of coarse classification based on visual classification and two-dimensional area thresholds, lacking deep coupling with defect depth, curvature, and material mechanical properties. This makes it difficult to directly provide downstream processes with physically meaningful quantitative grading guidance. Therefore, how to achieve high-precision three-dimensional reconstruction of multiple types of defects across the entire surface of an aluminum rod under the complex optical environment of high-speed rotation and strong specular reflection, accurately decouple similar defect features, and realize a collaborative mechanism between multimodal perception and the underlying mechanical model to complete a scientifically based quantitative grading of defects is a core technical problem that urgently needs to be solved.
[0005] To address this, a visual inspection and classification method for surface defects in aluminum rods is proposed. Summary of the Invention
[0006] The purpose of this invention is to provide a visual detection and classification method for surface defects on aluminum rods. It involves projecting multi-frequency sinusoidal structured light to acquire the original image sequence, performing sub-pixel-level vibration motion compensation, and generating a three-dimensional height map of the defect region through multi-frequency difference-frequency phase expansion and solution in cylindrical coordinates. A dual-stream Transformer network is used to fuse the two-dimensional grayscale image and the three-dimensional height map, and graph neural network inference is performed on the joint features to output visual classification results and confidence scores. When the confidence score is below a threshold, adaptive eddy current and ultrasonic detection are triggered. Multimodal anomaly features are aggregated through spatiotemporal registration and DS evidence theory to output defect type labels. Based on the three-dimensional height map, the equivalent volume, depth-to-diameter ratio, and minimum radius of curvature of the defect perimeter are calculated. Combined with the material compensation coefficient, the stress concentration factor is calculated, and the defect severity level is output. This invention overcomes the interference of high-speed rotation and strong reflectivity of the aluminum rod, improving the defect detection effect.
[0007] To achieve the above objectives, the present invention provides the following technical solution: A method for visual inspection and classification of surface defects in aluminum rods, comprising: The camera and projection parameters are determined based on the diameter of the aluminum rod. Multi-frequency sinusoidal structured light is projected and three-step phase-shift acquisition is performed to obtain the original image sequence containing deformation stripes. Vibration displacement signals are obtained using a three-axis Doppler vibration measurement device, and sub-pixel-level vibration motion compensation is performed on the original image sequence. Two-dimensional grayscale images and wrapping phase images of the compensated image sequence are extracted, and absolute phase reconstruction is completed by multi-frequency difference phase expansion; based on the structural cursor positioning parameters, physical size calculation is performed on the absolute phase in the cylindrical coordinate system of the aluminum rod to generate a three-dimensional height map of the defect area; A dual-stream Transformer network is used to encode the two-dimensional grayscale image and the three-dimensional height map through intensity stream and depth stream respectively, and outputs a joint feature vector through cross-modal cross-attention fusion. The joint feature vector is input into the aluminum rod defect knowledge graph to perform graph neural network inference, and the visual classification result and confidence score are output. When the confidence score is lower than the preset threshold, eddy current and ultrasonic detection are triggered. The multimodal anomaly features and visual classification results are spatially registered, temporally aligned and aggregated and fused to output the final defect type label. The equivalent volume, depth-to-diameter ratio, and minimum radius of curvature of the defect perimeter are calculated based on the three-dimensional height map. The stress concentration factor is calculated by combining the material compensation coefficient. A comprehensive quantitative score is constructed based on the equivalent volume and stress concentration factor, and the severity level of the defect is output according to the preset grading threshold.
[0008] Preferably, the step of determining camera and projection parameters based on the aluminum rod diameter, projecting multi-frequency sinusoidal structured light and performing three-step phase-shift acquisition to obtain the original image sequence containing deformation stripes includes: obtaining the diameter parameters of the aluminum rod to be inspected, and calculating the spatial chord length and maximum chord height of the corresponding arc surface of the aluminum rod in combination with a preset single detection center angle; determining the required field of view size based on the spatial chord length, and determining the required detection depth of field based on the maximum chord height, so as to solve for the camera and projection parameters, which include the camera working distance and the projection focal length; generating a sinusoidal structured light pattern with multiple spatial frequencies through a digital micromirror device, and performing three-step phase-shift projection on the surface of the aluminum rod at each spatial frequency, synchronously triggering the camera to acquire the reflected image modulated by the surface deformation, and obtaining the original image sequence containing deformation stripes.
[0009] Preferably, the subpixel-level vibration motion compensation includes: acquiring the three-dimensional real-time vibration displacement signal of the aluminum rod through a three-axis Doppler vibration measurement device, projecting it onto the imaging plane in combination with the camera working distance and projection focal length in the camera and projection parameters, and obtaining the inter-frame global spatial offset; using the first frame image as the position reference, performing pixel grayscale resampling mapping on each frame of the image sequence to be compensated based on the subpixel interpolation model, and outputting the motion-compensated image sequence.
[0010] Preferably, the multi-frequency difference phase unfolding to complete absolute phase reconstruction includes: performing three-step phase shift projection and acquisition on each of the three multi-frequency sinusoidal structured lights with spatial frequency decreasing relationship, and obtaining the wrapping phase map corresponding to each spatial frequency respectively; performing pixel-by-pixel difference frequency processing on the wrapping phase maps corresponding to adjacent spatial frequencies to generate a fundamental frequency phase map that covers the global field of view and has no phase folding; using the fundamental frequency phase map as the unfolding reference, the order of each frequency stripe is deduced step by step according to the frequency multiple relationship, and pixel-by-pixel stitching is performed in combination with the wrapping phase map of the highest spatial frequency to obtain a continuous absolute phase map.
[0011] Preferably, the step of performing physical dimension calculation on the absolute phase in the cylindrical coordinate system of the aluminum rod based on the structural light calibration parameters to generate a three-dimensional height map of the defect area includes: obtaining the structural light calibration parameters, which include the camera intrinsic parameter matrix, the projection device intrinsic parameter matrix, and the relative pose extrinsic parameters between the camera and the projection device; establishing a structured light triangulation geometric mapping model based on the camera and projection parameters and the absolute phase reconstruction results, and converting the phase information into spatial three-dimensional coordinates; extracting the absolute phase difference data and combining it with the reference phase constraint, mapping it to a cylindrical coordinate system with the aluminum rod axis as the central axis, and finally generating a three-dimensional height map of the defect area through coordinate transformation.
[0012] Preferably, the joint feature vector output through cross-modal cross-attention fusion includes: converting the two-dimensional grayscale image feature matrix output by intensity stream coding and the three-dimensional height map feature matrix output by depth stream coding into query, key, and value matrices through linear mapping, and performing bidirectional cross-modal cross-attention operation: using the two-dimensional grayscale image features as queries and the three-dimensional height map features as keys and values to obtain a depth-guided intensity attention representation; simultaneously using the three-dimensional height map features as queries and the two-dimensional grayscale image features as keys and values to obtain an intensity-guided depth attention representation; concatenating the two in the channel dimension and then performing nonlinear mapping through a multilayer perceptron to output a joint feature vector.
[0013] Preferably, the execution of graph neural network inference and the output of visual classification results and confidence scores include: mapping the joint feature vector as the initial embedding vector of the defect node to be tested to the aluminum rod defect knowledge graph, the knowledge graph containing defect instance nodes, defect type nodes, process cause nodes and their associated edges; after aggregating the features of nearest neighbor nodes through a graph attention network, calculating the probability distribution of each defect category through a fully connected classification layer Softmax, taking the category with the highest probability as the visual classification result, and taking the highest probability value as the confidence score.
[0014] Preferably, the spatial registration, temporal alignment, and aggregation fusion include: in the spatial registration dimension, using the global coordinate system of the visual acquisition device as a reference, the two-dimensional scanning coordinates of the eddy current probe and the depth coordinates of the ultrasonic probe obtained by eddy current and ultrasonic detection are uniformly mapped to the global coordinate system through rigid body transformation to obtain the spatial position registration results of multimodal anomaly features; in the temporal alignment dimension, using the visual acquisition timestamp as a reference, the pulse counting information on the aluminum rod conveying device is extracted, the acquisition delay of each mode is calculated, and compensation alignment is performed; in the aggregation fusion dimension, the probability distribution of the visual classification results, the eddy current anomaly impedance amplitude characteristics and ultrasonic echo attenuation rate characteristics corresponding to eddy current and ultrasonic detection are respectively converted into basic probability allocation functions, and multi-source confidence is aggregated through the orthogonal sum rule of DS evidence theory, and the proposition with the highest probability is taken as the final defect type label.
[0015] Preferably, the step of calculating the equivalent volume, depth-to-diameter ratio, and minimum radius of curvature of the defect perimeter based on the three-dimensional height map, and calculating the stress concentration factor in conjunction with the material compensation coefficient, includes: taking the surface height of the undeformed aluminum rod as zero reference, performing a two-dimensional spatial double integral on the absolute value of the height difference in the defect area to obtain the equivalent volume; extracting the maximum depth value of the defect area, converting the defect projection area into the equivalent circle diameter, and calculating the ratio of the maximum depth value to the equivalent circle diameter to obtain the depth-to-diameter ratio; extracting the three-dimensional surface point cloud data of the defect contour, calculating the principal curvature of each discrete point, and taking the reciprocal of the absolute value of the global maximum principal curvature as the minimum radius of curvature of the defect perimeter; and based on material mechanics constraints, constructing a three-dimensional input feature vector with the depth-to-diameter ratio, minimum radius of curvature, and material compensation coefficient, inputting it into a pre-calibrated stress concentration nonlinear regression model to perform mapping, and obtaining the stress concentration factor.
[0016] Preferably, the step of constructing a comprehensive quantitative score based on equivalent volume and stress concentration factor, and outputting the defect severity level according to a preset grading threshold, includes: performing a product operation on the normalized equivalent volume and stress concentration factor to obtain the basic damage base; performing an operation on the normalized stress concentration factor combined with the material sensitization coefficient to obtain the stress amplification coefficient; performing a linear weighted summation on the basic damage base and the stress amplification coefficient to output a comprehensive quantitative score; and comparing the comprehensive quantitative score with the preset grading threshold within an interval to determine and output the defect severity level.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention addresses the challenges of high-speed rotation and specular reflection in aluminum rods by introducing triaxial Doppler vibration displacement signals to perform subpixel-level motion compensation. Combined with multi-frequency sinusoidal structured light and difference-frequency phase unfolding technology, it effectively avoids stripe distortion and phase folding caused by high-speed motion. In complex optical environments, a high-resolution, distortion-free 3D height map of the defect region is successfully constructed, laying a high-quality data foundation for subsequent accurate evaluation.
[0018] 2. This invention employs a dual-stream Transformer network, utilizing intensity flow (two-dimensional grayscale image) and depth flow (three-dimensional height map) to perform bidirectional cross-modal cross-attention fusion, and combines it with an aluminum rod defect knowledge graph for graph neural network inference. This feature achieves precise decoupling of surface texture and deep semantics of three-dimensional morphology, enabling it to keenly capture the microscopic differences of highly deceptive defects such as pores and indentations, significantly improving classification confidence.
[0019] 3. This invention overcomes the detection blind spots of pure vision solutions by designing a dynamic detection path triggered by a confidence threshold: when the visual classification confidence is low, the system adaptively triggers eddy current and ultrasonic physical detection, and uses the orthogonal sum rules of DS evidence theory to spatially register and aggregate the visual probability distribution with the physical detection anomaly features. This multi-source information fusion mechanism compensates for the shortcomings of a single sensor.
[0020] 4. This invention combines geometric parameters such as equivalent volume, aspect ratio, and minimum radius of curvature of the defect perimeter calculated from a three-dimensional height diagram with specific aluminum alloy material compensation coefficients to deeply calculate the "stress concentration factor" of defects. By constructing a comprehensive quantitative score that includes the basic damage base and stress amplification factor, this method synergistically analyzes surface detection and underlying fatigue fracture mechanics mechanisms, providing support for downstream production process downgrading or release decisions. Attached Figure Description
[0021] Fig. 1 This is a flowchart illustrating a method for visual inspection and classification of surface defects in aluminum rods, provided in an embodiment of the present invention. Fig. 2 This is a schematic diagram illustrating the generation of a three-dimensional height map of a defect region, provided as an embodiment of the present invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] Please see Figs. 1-2 This invention provides a method for visual inspection and classification of surface defects in aluminum rods, the technical solution of which is as follows: Example 1: This embodiment uses the defect detection of aluminum bars in an aluminum processing plant as a specific application scenario, and applies a visual inspection and classification method for surface defects of aluminum bars, including: The camera and projection parameters are determined based on the diameter of the aluminum rod. Multi-frequency sinusoidal structured light is projected and three-step phase-shift acquisition is performed to obtain the original image sequence containing deformation stripes. Vibration displacement signals are obtained using a three-axis Doppler vibration measurement device, and sub-pixel-level vibration motion compensation is performed on the original image sequence. Two-dimensional grayscale images and wrapping phase images of the compensated image sequence are extracted, and absolute phase reconstruction is completed by multi-frequency difference phase expansion; based on the structural cursor positioning parameters, physical size calculation is performed on the absolute phase in the cylindrical coordinate system of the aluminum rod to generate a three-dimensional height map of the defect area; A dual-stream Transformer network is used to encode the two-dimensional grayscale image and the three-dimensional height map through intensity stream and depth stream respectively, and outputs a joint feature vector through cross-modal cross-attention fusion. The joint feature vector is input into the aluminum rod defect knowledge graph to perform graph neural network inference, and the visual classification result and confidence score are output. When the confidence score is lower than the preset threshold, eddy current and ultrasonic detection are triggered. The multimodal anomaly features and visual classification results are spatially registered, temporally aligned and aggregated and fused to output the final defect type label. The equivalent volume, depth-to-diameter ratio, and minimum radius of curvature of the defect perimeter are calculated based on the three-dimensional height map. The stress concentration factor is calculated by combining the material compensation coefficient. A comprehensive quantitative score is constructed based on the equivalent volume and stress concentration factor, and the severity level of the defect is output according to the preset grading threshold.
[0024] Further, the step of determining camera and projection parameters based on the aluminum rod diameter, projecting multi-frequency sinusoidal structured light and performing three-step phase-shift acquisition to obtain the original image sequence containing deformation stripes includes: obtaining the diameter parameters of the aluminum rod to be inspected, and calculating the spatial chord length and maximum chord height of the corresponding arc surface of the aluminum rod in combination with a preset single detection center angle; determining the required field of view size based on the spatial chord length, and determining the required detection depth of field based on the maximum chord height, so as to solve for the camera and projection parameters, which include the camera working distance and the projection focal length; generating sinusoidal structured light patterns of multiple spatial frequencies through a digital micromirror device, and performing three-step phase-shift projection on the surface of the aluminum rod at each spatial frequency, synchronously triggering the camera to acquire the reflected image modulated by the surface deformation, and obtaining the original image sequence containing deformation stripes.
[0025] Specifically, in this embodiment, the nominal diameter D of the aluminum rod to be inspected is 100mm, and the central angle θ of a single inspection is preset to 60°. With the axis of the aluminum rod as the center, the chord length L corresponding to the arc surface to be inspected is determined by geometric relationships: L = 2·(D / 2)·sin(θ / 2) = 2×50×sin30° = 50mm; the corresponding chord height h = (D / 2)·(1 - cos(θ / 2)) = 50×(1 - cos30°)≈ 6.70mm. The field of view (FOV) is set to be no less than the chord length L, i.e., FOV ≥ 50mm. In this embodiment, FOV = 55mm is used to cover the edge of the arc surface with a margin. The depth of field (DOF) is set to be no less than 1.5 times the maximum chord height h, i.e., DOF ≥ 10.05mm. In this embodiment, DOF = 12mm is used to ensure that surface details at all positions on the arc surface are clearly imaged within the depth of field.
[0026] Based on the aforementioned field of view and depth of field constraints, and combined with the industrial camera imaging formula, the calculated working distance WD (distance from the principal point in front of the camera lens to the apex of the arc surface of the aluminum rod) is 300mm. An industrial lens with a focal length of f=25mm and an industrial area array camera with a resolution of 2448×2048 pixels and a pixel size of 3.45μm are selected, corresponding to a single-pixel projection resolution of approximately 22.5μm, which meets the defect detection requirement of a minimum width of not less than 0.05 mm. The projector focal length is determined synchronously based on the projection distance and the structured light fringe pitch requirement: in this embodiment, the projection focal length is 15mm, which, in conjunction with a digital micromirror device (DMD, resolution 1920×1080, pixel size 7.56μm), achieves the highest spatial frequency. =16 lines / mm, medium frequency =4 lines / mm and low frequency =1 line / mm for precise projection of sinusoidal structured light patterns at three spatial frequencies. The frequency ratios between the three frequencies satisfy... The integer multiple relationship of 16:4:1 provides a frequency multiple basis for subsequent multi-frequency difference frequency phase expansion.
[0027] The three-step phase-shift projection process is as follows: For each spatial frequency... (k∈{l,m,h}), three sinusoidal fringe patterns with phase shifts of 0°, 120°, and 240° are sequentially loaded onto the surface of an aluminum rod by a DMD, and projected one by one. During each projection, the camera synchronizes with the DMD pattern refresh via a hardware trigger signal, and the exposure time is set to 500μs. With a linear velocity of 1.5 m / s for the aluminum rod, the axial displacement during a single-frame exposure is approximately 0.75 mm. Therefore, the original acquired image contains fringe deformation caused by motion, which is corrected through subsequent vibration compensation and phase calculation. After the three-step phase-shift acquisition is completed, three original grayscale images containing surface deformation modulation fringes are obtained for each frequency. (n=1,2,3), the expression for the grayscale value of the nth image at the kth spatial frequency is: = + ·cos[ + (n-1)×2 / 3], where A(x,y) is the background light intensity, B(x,y) is the stripe contrast modulation amplitude, The phase value is the wrapping phase value modulated by surface deformation, and (x,y) are the pixel coordinates. At this point, the acquisition of the original image sequence is complete, yielding a total of 3×3=9 original images, which constitute the input data for subsequent vibration compensation and phase calculation.
[0028] Furthermore, the subpixel-level vibration motion compensation includes: acquiring the three-dimensional real-time vibration displacement signal of the aluminum rod through a three-axis Doppler vibration measurement device, projecting it onto the imaging plane in combination with the camera working distance and projection focal length in the camera and projection parameters, and obtaining the inter-frame global spatial offset; using the first frame image as the position reference, performing pixel grayscale resampling mapping on each frame of the image sequence to be compensated based on the subpixel interpolation model, and outputting the motion-compensated image sequence.
[0029] Specifically, a triaxial Doppler laser vibrometer (LDV, model similar to Polytec PSV-500, sampling frequency 100kHz) is positioned to the side of the aluminum rod inspection station. Its three laser beams illuminate the surface of the aluminum rod along the radial (Y-axis), axial (Z-axis), and imaging optical axis (X-axis) directions, respectively, to acquire vibration velocity signals in real time in these three directions. , , By integrating the velocity signals over time, a three-dimensional vibration displacement synchronized with the camera trigger moment is obtained. , , (Unit: μm)
[0030] The three-dimensional vibration displacement is projected onto the camera imaging plane to obtain the global pixel offset between frames. , ): Horizontal offset within the imaging plane = Vertical offset = ,in WD represents the camera lens focal length (25mm) and the working distance (300mm). The camera pixel size is 3.45 μm. The above formula maps vibrational displacement in physical space to sub-pixel offset in pixel units through perspective projection. Displacement along the optical axis... Under the influence of [unspecified factors], a slight change in the imaging scale will occur. In this embodiment, when the working distance is 300 mm and the typical vibration amplitude does not exceed 50 μm, the introduced scale change is less than 0.017%, and the impact on pixel translation is negligible. Therefore, only [unspecified method] is used. and Component calculations determine pixel offsets within the image plane.
[0031] Using the first frame as the position reference (corresponding to the camera trigger time) ), for subsequent frames of images (n≥2) Perform reverse translation compensation: adjust the offset of each frame relative to the first frame ( The original image is inverted, and a bicubic subpixel interpolation model is used to resample and map the grayscale values, outputting the compensated images for each frame. The bicubic interpolation extracts 16 pixels from a 4×4 neighborhood, and the grayscale values at the target coordinates are reconstructed using cubic polynomial weighted summation. The interpolation accuracy is better than 0.1 pixels, meeting the pixel-level alignment accuracy requirement corresponding to a minimum defect width of approximately 0.05 mm. Through the above compensation process, the stripe patterns in each frame of the compensated image sequence are aligned to the position of the aluminum rod in the first frame, providing a consistent image input for subsequent phase calculations.
[0032] Furthermore, the multi-frequency difference phase unfolding to complete absolute phase reconstruction includes: performing three-step phase shift projection and acquisition on each of the three multi-frequency sinusoidal structured lights with spatial frequency decreasing relationship, and obtaining the wrapping phase map corresponding to each spatial frequency respectively; performing pixel-by-pixel difference frequency processing on the wrapping phase maps corresponding to adjacent spatial frequencies to generate a fundamental frequency phase map that covers the global field of view and has no phase folding; using the fundamental frequency phase map as the unfolding reference, the order of each frequency stripe is deduced step by step according to the frequency multiple relationship, and pixel-by-pixel stitching is performed in combination with the wrapping phase map of the highest spatial frequency to obtain a continuous absolute phase map.
[0033] Specifically, for the vibration-compensated image sequence, for each spatial frequency... (Unit: strips / mm), the wrapping phase map is extracted pixel by pixel using the three-step phase shift formula. The wrapping phase is calculated as follows: The resulting package phase range is (- , ], corresponding to the phase folding range of each frequency.
[0034] Differential frequency processing steps: First, process the high-frequency wrapping phase diagram. Phase diagram with intermediate frequency wrap Perform differential frequency per pixel: ,in [·] represents the phase winding operator, and the specific formula is as follows: Used to constrain the phase result to (- , The equivalent spatial frequency after the difference frequency is: = 12 lines / mm. Further analysis of the difference frequency results. Low-frequency wrapping phase diagram Perform the difference frequency calculation again: The corresponding equivalent spatial frequency is: =11 strips / mm.
[0035] Based on the structured light triangulation principle, the system's phase-free folding depth measurement range, fringe spatial frequency, and system geometric parameters (including the projection-camera baseline distance) are considered. and working distance Related to this, the baseline distance generally increases as the spatial frequency decreases. In this embodiment, the baseline distance... =120mm, camera working distance =300mm, corresponding to the equivalent low frequency The unambiguous depth measurement range of 11 lines / mm is much larger than the maximum defect depth that may occur on the aluminum rod surface (2 mm). Therefore, in other words, the absolute phase changes caused within this depth range are all within (- , Within the interval, no phase folding will occur, and the resulting fundamental frequency phase... It can serve as an absolute reference for multi-frequency phase expansion.
[0036] The absolute phase deduction process is as follows: using the fundamental frequency phase diagram... As a reference for expansion, the fringe order of the mid-frequency wrapping phase is inversely calculated. Specifically: ; in {·} represents the nearest integer rounding operation.
[0037] This leads to the intermediate frequency absolute phase. In the same way, using the intermediate frequency absolute phase... Inversely calculate the order of high-frequency stripes And obtain high-frequency absolute phase , The high-frequency absolute phase is a continuous and unambiguous phase distribution, and its spatial resolution is determined by the highest frequency. The 16 stripes / mm ratio corresponds to a stripe pitch of approximately 62.5μm, which determines the accuracy of subsequent 3D height reconstruction.
[0038] Furthermore, the step of performing physical dimension calculation on the absolute phase in the cylindrical coordinate system of the aluminum rod based on the structural beam positioning parameters to generate a three-dimensional height map of the defect area includes: obtaining the structural beam positioning parameters, which include the camera intrinsic parameter matrix, the projection device intrinsic parameter matrix, and the relative pose extrinsic parameters between the camera and the projection device; establishing a structured light triangulation geometric mapping model based on the camera and projection parameters and the absolute phase reconstruction results, converting the phase information into spatial three-dimensional coordinates; extracting the absolute phase difference data and combining it with the reference phase constraint, mapping it to a cylindrical coordinate system with the aluminum rod axis as the central axis, and finally generating a three-dimensional height map of the defect area through coordinate transformation.
[0039] Specifically, the structural calibration employs the standard planar target method (target grid spacing accuracy ≤ 1μm), performing intrinsic parameter calibration on both the camera and the projection device to obtain the camera intrinsic parameter matrix. (Including focal length) , Main point ( , (and distortion coefficients) and the intrinsic parameter matrix of the projection device The rotation matrix between the camera and the projection device is obtained through joint calibration. and translation vector (Collectively referred to as external parameters). Baseline distance (The physical distance between the optical center of the camera and the optical center of the projection device) is determined by the translation vector. The modulus length is determined in this embodiment. =120mm.
[0040] Based on the above calibration parameters, a geometric model for structured light triangulation is established. Under the conditions of pinhole imaging model and small-angle approximation, the spatial points on the aluminum rod surface are... Depth value in camera coordinate system (i.e., point) Distance along the camera's optical axis and absolute phase difference The relationship between them can be approximated as: = × × / ( ) in, Baseline distance; The equivalent focal length of the projection device; The absolute phase difference; For high-frequency spatial frequencies (bars / mm). This refers to the calibration scale parameter from the reference plane to the projection reference coordinate system (determined during the system calibration phase, used to establish the proportional relationship between phase and spatial distance). This relationship is based on the correspondence between the structured light phase and the spatial position of the projected fringes, mapping phase information to spatial depth information. The absolute phase difference is defined as: = - ,in, The high-frequency absolute phase obtained from the aforementioned multi-frequency phase expansion. This is a reference phase map pre-acquired under the condition of a defect-free ideal cylindrical surface, which is defined in the same pixel coordinate system as the current phase.
[0041] After obtaining the depth information, the three-dimensional coordinates of each pixel in the camera coordinate system are... Through camera intrinsic parameter matrix Inverse mapping and extrinsic parameters Perform a coordinate transformation, mapping to a cylindrical coordinate system centered on the axis of the aluminum rod. The cylindrical coordinate system is represented as follows: ,in Circumferential angle, This refers to the axial position. This represents the radial distance.
[0042] Based on the nominal radius of the aluminum rod Using mm=50mm as the reference radius, the height deviation (i.e., defect depth or protrusion height) of each pixel is defined as follows: Perform the above coordinate transformation and height calculation on all pixels within the detection area to construct a coordinate system based on the circumferential angle. and axial position A three-dimensional height map of the defect region is generated, with the coordinate axes and height deviation as pixel values. The spatial resolution of the height map is better than 25μm×25μm, and the height measurement resolution is better than 1μm, which can meet the requirements for accurate characterization of small defects on the surface of aluminum rods.
[0043] Fig. 2 This is a schematic diagram illustrating the generation of a three-dimensional height map of a defect region, provided as an embodiment of the present invention.
[0044] Furthermore, the joint feature vector output through cross-modal cross-attention fusion includes: converting the two-dimensional grayscale image feature matrix output by intensity stream coding and the three-dimensional height map feature matrix output by depth stream coding into query, key, and value matrices through linear mapping, and performing bidirectional cross-modal cross-attention operation: using the two-dimensional grayscale image features as queries and the three-dimensional height map features as keys and values to obtain a depth-guided intensity attention representation; simultaneously using the three-dimensional height map features as queries and the two-dimensional grayscale image features as keys and values to obtain an intensity-guided depth attention representation; concatenating the two in the channel dimension and then performing nonlinear mapping through a multilayer perceptron to output a joint feature vector.
[0045] Specifically, the dual-stream Transformer network consists of two parallel encoding branches, intensity stream and depth stream, and a cross-modal cross-attention fusion module. The intensity stream takes a 2D grayscale image (448×448 pixels, single channel) as input, while the depth stream takes a 3D height map (448×448 pixels, single channel, normalized to the [-1,1] interval) as input. The original 2448×2048 pixel structured light image and its corresponding height map are post-processed after a defect candidate region (ROI) detection step: the depth deviation in the 3D height map exceeds a deviation threshold. Centered on a 5μm connected region centroid, a minimum bounding rectangle region of 20 pixels is extracted and scaled to 448×448 pixels using bilinear interpolation, serving as the input for a two-stream network. ROI detection ensures that the target defect constitutes at least 30% of the pixels in the 448×448 input image to avoid losing crucial texture information of minute defects with a width ≥0.05mm (approximately 2.2 pixels) during downsampling. For cases with multiple separate defect regions within a single frame, ROIs are extracted and inferred independently for each defect.
[0046] Deviation threshold =5μm, the setting is based on: extracting the three sigma upper limit of the normal surface height deviation by performing structured light measurement on the surface of a defect-free reference aluminum rod. And introduce an adaptive safety factor Perform calculations, that is = × Taking a 6061-T6 aluminum rod with a surface roughness Ra≈1.2μm as an example, its statistically obtained... Approximately 3μm (obtained from the statistical standard deviation of the height distribution of a defect-free reference surface). In this embodiment, k=1.5 is used, therefore, it is set as follows: =5μm. For aluminum bars of different grades and extrusion processes, this threshold can be linearly adaptively calibrated according to the change of the reference surface roughness to ensure that the interference of normal surface roughness is eliminated within the three sigma range; Both branches employ the Vision Transformer (ViT-B / 16) backbone network: the input image is divided into 16×16 pixel image patches. Each patch is linearly embedded and projected into a token vector of dimension d=768. This token vector is then added to a learnable positional encoding and input into a 12-layer Transformer encoder (each layer contains an 8-head self-attention mechanism and a feedforward neural network, with a hidden layer dimension of 3072). The positional encoding is used to preserve the spatial structure information of the image, ensuring that the tokens retain their spatial correspondence during cross-modal attention computation.
[0047] Intensity flow and depth flow each output two-dimensional grayscale image feature matrices: With the 3D height map feature matrix Where N = (448 / 16) 2=784 represents the number of tokens, and d=768 represents the feature dimension. The ViT-B / 16 backbone network weights of the two branches are initialized using ImageNet-21k pre-trained weights and fine-tuned end-to-end on an aluminum rod defect dataset (containing 5 types of defects: porosity, indentation, inclusion, oxidation spalling, and scratches, with ≥2000 samples per type, divided into training / validation / test sets in a 7:2:1 ratio). The learning rate is 1e-4, using the AdamW optimizer, and training for 50 epochs. The training dataset was collected from the online detection system of aluminum rod manufacturer A (main product specifications D=60~150mm), containing 2150 samples of 5 types of defects: 2300 for porosity, 2080 for inclusion, 2200 for oxidation spalling, and 2100 for scratches, with ≥2000 samples of each type, and 5000 samples without defects, with a positive to negative sample ratio of approximately 2:1. Samples for each category were collected under different aluminum rod diameters (D=60mm, 100mm, and 150mm), different linear velocities (1.0m / s, 1.5m / s, and 2.0m / s), and different ambient light intensities to ensure data diversity. Each pair of structured light images and height maps was verified by annotators to ensure consistent pairing and correspondence with the same timestamp. Data augmentation strategies included random cropping (±20 pixels), horizontal flipping, brightness jitter (±10%), and Gaussian noise injection. =0.01). The training set, validation set, and test set are divided in a 7:2:1 ratio, and stratified sampling ensures that the proportion of each category is consistent in each subset.
[0048] In the cross-modal attention fusion module, and Each is passed through three independent sets of learnable linear projection matrices , , (All dimensions are) × , =64) is mapped to a query matrix Key matrix AND-value matrix .
[0049] First direction (deeply guiding intense attention): = · , = · , = · Calculate the attention output AID = Softmax( · / )· Obtain the guided attention representation of the depth mode on the intensity mode AID∈ .
[0050] Second direction (intensity-guided deep attention): = · , = · , = · Similarly, the intensity-guided deep attention representation ADI∈ is obtained. Concatenate AID and ADI along the channel (feature) dimension to obtain the concatenated feature [AID; ADI]∈ The concatenated features are input into a two-layer fully connected (MLP) network (512 hidden layers, ReLU activation function, 0.1 dropout rate), and compressed to 1×2 using global average pooling. =128-dimensional vector, output is joint feature vector ∈ This vector simultaneously encodes cross-modal association information between surface grayscale texture features and three-dimensional shape features, which is used for subsequent knowledge graph reasoning.
[0051] Furthermore, the execution of graph neural network inference and the output of visual classification results and confidence scores include: mapping the joint feature vector as the initial embedding vector of the defect node to be tested to the aluminum rod defect knowledge graph, the knowledge graph containing defect instance nodes, defect type nodes, process cause nodes and their associated edges; after aggregating the features of nearest neighbor nodes through a graph attention network, calculating the probability distribution of each defect category through a fully connected classification layer Softmax, taking the category with the highest probability as the visual classification result, and taking the highest probability value as the confidence score.
[0052] Specifically, the node set V of the aluminum rod defect knowledge graph G=(V,E) consists of three types of nodes: Defect instance node vinst: Historical defect samples stored in the graph. Each instance node carries a 128-dimensional feature embedding vector (consistent with the output format of the two-stream Transformer) and its corresponding label. The defect type node vtype includes five defect category nodes: porosity, indentation, inclusion, oxidation spalling, and scratch. Each node carries a prototype feature vector for each defect type (obtained by aggregating the average features of all instance nodes of the same category). Specifically, the underlying sensor process parameters of each process (including melting and casting temperature, extrusion speed, and cooling water pressure) are extracted, processed by max-min normalization, and then concatenated into a process parameter vector. This vector is then mapped to a 128-dimensional process feature description vector through a single-layer fully connected network to ensure alignment with the feature dimensions of the defect instance nodes and to meet the feature aggregation requirements of the subsequent graph attention network.
[0053] Process cause node vcause: Aluminum bar production process node (melting, casting, homogenization, extrusion, cooling, etc.), carrying process feature description vectors. Edge set E includes: "belonging" relationship edges between instance nodes and their corresponding type nodes, "cause association" relationship edges between type nodes and process cause nodes (constructed offline based on process knowledge), and "similar" relationship edges between instance nodes based on feature cosine similarity (edges are established when similarity > 0.85). Online graph maintenance strategy: Each time a new high-confidence detection sample is added, its instance node and corresponding relationship edges are appended to the graph to keep the graph dynamically expanding. Dynamic graph expansion adopts a batch update strategy: After accumulating 50 new instance nodes, the feature vector of the new node is used as the query to search for historical nodes with cosine similarity > 0.85 in the entire graph, and 'similar' relationship edges are established or updated in batches; a hierarchical navigable small world graph (HNSW) index is used to accelerate nearest neighbor retrieval. When the graph size reaches 100,000 nodes, the Top-10 retrieval latency is still less than 5ms (M=32, ef_construction=200). If an instance node is not hit by any query in 200 consecutive inferences, it is marked as an inactive node and removed from the online retrieval index (the node data is retained for offline analysis) to control the size of the online graph and suppress the impact of noise sample accumulation on classification accuracy.
[0054] Joint feature vectors As the initial embedding vector of the defect node vquery to be tested, the Top-K (K=10) instance nodes with the highest feature similarity to vquery are retrieved in the graph. A local subgraph is constructed with vquery as the center node, K nearest neighbor instance nodes, and their associated type nodes and causal nodes. Graph Attention Network (GAT) inference is performed on this local subgraph. GAT has 3 layers. When aggregating the features of nearest neighbors in each layer, a multi-head attention mechanism (8 heads, attention dimension 16) is used. The attention coefficient is calculated by the dot product similarity between the embedding vector of the center node and the embedding vectors of the nearest neighbors. After Softmax normalization, the nearest neighbor features are weighted and aggregated to update the representation of the center node. The updated embedding vector hquery of the center node output by the 3rd layer of GAT is 128-dimensional. After passing through a fully connected layer (128→5) and the Softmax activation function, the probability distribution P={p1, p2, p3, p4, p5} of 5 defect categories is output. The category corresponding to the highest probability is taken as the visual classification result, and the highest probability value pmax=max(P) is taken as the confidence level. (Preset threshold) Based on ROC curve analysis on the validation set, the maximum confidence cutoff point is selected under the constraint of a false alarm rate <2%. In this embodiment... ):when When, directly output the visual classification result; when At that time, the eddy current and ultrasonic multimodal fusion verification process is triggered. When the confidence level is greater than or equal to a preset threshold, the visual classification result is directly output as the final defect type label.
[0055] In terms of hardware deployment, to match the linear velocity of 1.5 m / s and the detection field of view of 55 mm, the system main control adopts a multi-threaded asynchronous pipeline architecture. A single image acquisition takes approximately 30 ms (with a 45 mm displacement during this period to ensure image overlap), and while a single visual algorithm inference takes approximately 120 ms, the parallel computing across 4 nodes ensures that the overall system throughput seamlessly matches the acquisition frame rate, achieving blind-zone-free detection. Furthermore, the eddy current and ultrasonic probes are asynchronously triggered by the lower-level machine and executed at downstream workstations, without obstructing the main visual inspection pipeline.
[0056] Furthermore, the spatial registration, temporal alignment, and aggregation fusion include: in the spatial registration dimension, using the global coordinate system of the visual acquisition device as a reference, the two-dimensional scanning coordinates of the eddy current probe and the depth coordinates of the ultrasonic probe obtained by eddy current and ultrasonic detection are uniformly mapped to the global coordinate system through rigid body transformation to obtain the spatial position registration results of multimodal anomaly features; in the temporal alignment dimension, using the visual acquisition timestamp as a reference, the pulse counting information on the aluminum rod conveying device is extracted, the acquisition delay of each mode is calculated, and compensation alignment is performed; in the aggregation fusion dimension, the probability distribution of the visual classification results, the eddy current anomaly impedance amplitude characteristics corresponding to eddy current and ultrasonic detection, and the ultrasonic echo attenuation rate characteristics are respectively converted into basic probability allocation functions, and multi-source confidence is aggregated through the orthogonal sum rule of DS evidence theory, and the proposition with the highest probability is taken as the final defect type label.
[0057] Specifically, the spatial registration of the multimodal detection system adopts a unified global coordinate system strategy: the image coordinate system of the vision camera (already aligned with the cylindrical coordinate system of the aluminum rod through system calibration) serves as the global coordinate system reference. The eddy current probe (coil diameter 3mm, operating frequency 100kHz~1MHz) scans along the axis of the aluminum rod in a helical trajectory, and its two-dimensional scanning coordinates (axial feed position) are... Azimuth Combined with the cross-sectional radius of the aluminum rod being measured. Convert the azimuth angle to the circumferential arc length. Then, through the rigid pose relationship (rotation matrix) installed on the same testing platform, Translation vector (Obtained through offline calibration using a multimodal calibration board, calibration accuracy ≤ 0.1mm) and mapped to the global coordinate system via rigid body transformation: [ ; ] = ·[ ; ] + Then the mapped global arc length Convert back to global azimuth To complete the spatial registration in cylindrical coordinates.
[0058] An ultrasound probe (center frequency 5MHz, focal length 30mm) provides coordinates of depth-direction (radial) anomalies. (Based on ultrasound echo time of flight) With the speed of sound in an aluminum rod =6300m / s Calculation: s= × / 2), which is also mapped to the global coordinate system through rigid body transformation, and achieves sub-millimeter level (≤0.3mm) position registration with eddy current and visual inspection results in three-dimensional space.
[0059] Timing alignment employs a pulse encoder reference strategy: an incremental rotary encoder with a resolution of 1000 pulses / revolution is installed on the aluminum bar conveyor rollers. The roller diameter (Droller) is 100mm, corresponding to an axial linear resolution of approximately 0.314mm / pulse. Using the pulse count Nvis corresponding to the visual trigger moment as a reference, and the pulse count Neddy corresponding to the eddy current scan trigger moment, the axial displacement compensation amount of the eddy current relative to the visual perception is... × 0.314mm. The ultrasonic detection time delay is calculated similarly to ensure that the three-modal data correspond to the detection results at the same position on the aluminum rod in the time dimension, and the time alignment accuracy is better than 1 pulse count (approximately 0.3mm axial displacement).
[0060] The DS evidence theory aggregation and fusion process is as follows: the detection results of the three modalities are each converted into a recognition framework. The Basic Probability Assignment (BPA) function is used for the following conditions: pores, indentations, inclusions, oxidation spalling, and scratches. The visual modality BPA directly takes the probability distribution from the output of the Softmax graph neural network, letting mvis({ci}) = pi (i=1,...,5). (Assigned to the entire set without uncertainty).
[0061] Eddy current mode: based on the amplitude of eddy current abnormal impedance | | (unit: Ω) is the input, and a single-layer fully connected network trained offline (trained with 4000 labeled eddy current signal data, with | as the input) is used. | with phase angle The concatenated 2D feature vectors (the output layer contains 5 nodes and uses the Softmax activation function to calculate the probability distribution of each independent class) are converted into a BPA function. .
[0062] Ultrasonic Modal Analysis: The ultrasonic modal analysis uses a four-dimensional feature vector composed of ultrasonic echo attenuation rate α (dB / mm, calculated by dividing the difference in amplitude between the bottom echo and the defect echo by the propagation distance), bottom echo amplitude Abottom (dB), the amplitude ratio of the defect echo to the bottom echo Aratio (dimensionless), and ultrasonic center frequency offset Δf (kHz, calculated by the difference between the centroid of the defect echo spectrum and the reference frequency). The input is a multilayer perceptron (MLP) network containing two hidden layers (32 and 64 nodes respectively, using the ReLU activation function) and one output layer (5 nodes, using the Softmax activation function). α reflects the overall scattering intensity of the defect, Abottom reflects the ultrasonic penetration capability, Aratio distinguishes between surface and volumetric defects, and Δf reflects the defect particle size distribution characteristics. The combined use of these four features can effectively reduce the classification confusion rate of similar defect types such as pores and inclusions, and oxidation exfoliation and inclusions. This network was trained offline using a pre-collected dataset of 2000 sets of measured ultrasonic signals, and the output probability distribution of the five defect types was directly converted into the BPA function mus. The DS evidence theory orthogonal sum rule (pairwise aggregation) is applied sequentially to the three BPA functions: mfused = mvis ⊕ meddy ⊕ mus, where ⊕ represents the DS orthogonal sum operation (including conflict normalization processing; when the conflict coefficient K exceeds 0.9, the Yager correction rule is activated to improve robustness). When the confidence of the most probable proposition in the DS aggregation result is lower than the decision threshold δ=0.4 (i.e., the three-modal evidence conflict is too large, resulting in high uncertainty in the aggregation result), the system automatically activates a downgrade decision mechanism: the single modality output with the highest confidence is prioritized as the final defect type label, and the detection interface is marked 'High conflict exists in multi-source fusion (K>0.9), the result is for reference only, manual review is recommended'. At the same time, the original three-modal feature vectors and conflict coefficient K of this batch are recorded to the anomaly log for subsequent offline optimization iteration of the evidence network. In the actual scenario, according to the multimodal detection statistics of 1000 sample aluminum rods, the high conflict rate of K>0.9 is about 2.3%, mainly concentrated in the mixed defect scenario of oxidation spalling and slag inclusion. Finally, the defect category corresponding to the proposition with the highest probability in mfused is selected as the final defect type label for output.
[0063] Furthermore, the calculation of the equivalent volume, depth-to-diameter ratio, and minimum radius of curvature of the defect perimeter based on the three-dimensional height map, combined with the material compensation coefficient to calculate the stress concentration factor, includes: taking the surface height of the undeformed aluminum rod as zero reference, performing a two-dimensional spatial double integral on the absolute value of the height difference in the defect area to obtain the equivalent volume; extracting the maximum depth value of the defect area, converting the defect projection area into the equivalent circle diameter, and calculating the ratio of the maximum depth value to the equivalent circle diameter to obtain the depth-to-diameter ratio; extracting the three-dimensional surface point cloud data of the defect contour, calculating the principal curvature of each discrete point, and taking the reciprocal of the absolute value of the global maximum principal curvature as the minimum radius of curvature of the defect perimeter; based on material mechanics constraints, constructing a three-dimensional input feature vector with the depth-to-diameter ratio, minimum radius of curvature, and material compensation coefficient, inputting it into a pre-calibrated stress concentration nonlinear regression model to perform mapping, and obtaining the stress concentration factor. The stress concentration factor is an engineering equivalent parameter used to characterize the degree of local stress amplification in defects. Specifically, equivalent volume Calculation: The height of the normal aluminum rod surface outside the defect area (i.e., Using the 0 datum plane as a reference, perform a numerical double integral on all pixels belonging to the defect region in the 3D height map: ; in =25μm is the actual pixel size (determined by the structural cursor parameters), and the summation range is all pixels (i,j) within the defect area. During calculation... and Convert to mm.
[0064] Calculation of depth-to-diameter ratio γ: Extract the maximum depth value from the defect area height map. Unit: mm; Calculate the projected area of the defect. ( (Total number of pixels in the defect area), converted to the equivalent circle diameter. De = 2× The depth-to-diameter ratio is defined as γ= / (Dimensionless parameter) This parameter reflects the elongation of the defect shape. The larger the depth-to-diameter ratio, the sharper the defect and the greater its contribution to stress concentration.
[0065] Minimum radius of curvature of the defect perimeter Calculation: Extract the 3D surface point cloud (coordinates of points on the iso-depth contour line) from the 3D height map on the defect contour boundary, and calculate the principal curvature for each discrete boundary point. , (Obtained by fitting a local quadratic surface and then finding the eigenvalues of the Hessian matrix), the absolute value of the maximum principal curvature at this point is defined as... = max(| ,|,| |), corresponding to the local minimum radius of curvature. Traverse all discrete points on the contour and take the global maximum principal curvature: (Unit: mm). The minimum radius of curvature of the defect perimeter is defined as... (Unit: mm) The smaller the value, the sharper the defect profile and the more severe the theoretical stress concentration.
[0066] Stress Concentration Factor Calculation of material compensation coefficient Related to the alloy grade of aluminum rods, determined by the material's mechanical properties (yield strength) The normalized derivation of the elastic modulus E is based on 6061-T6 aluminum alloy. =276MPa, E=68.9GPa, =1.0); 6063-T5 aluminum alloy ( =145MPa, E=68.9GPa, =0.85); 7075-T6 aluminum alloy ( =503MPa, E=71.7GPa, =1.28). The aspect ratio, minimum radius of curvature, and material compensation coefficient are used to construct the 3D input feature vector (γ, , Input a pre-calibrated nonlinear regression model of stress concentration; the model is trained using a finite element simulation dataset, which covers γ∈[0.01,0.5]. The parameter space is ∈ [0.05, 2.0] mm, with a total of 3000 simulation conditions. Support Vector Regression (SVR) with an RBF kernel is used, and the parameter is set to C=100. Perform fitting and output stress concentration factor It is dimensionless, typically ranging from 1.0 to 4.5. In the three-dimensional input feature vector of the SVR model, The training coverage is [0.80, 1.35] (step size 0.05, including three validation points: 0.85 for 6063-T5, 1.00 for 6061-T6, and 1.28 for 7075-T6), and is related to γ and 3000 combined operating conditions were uniformly sampled within the parameter space. Five-fold cross-validation was used to evaluate the model performance: the coefficient of determination R0 on the reserved test set (600 cases) was used. 2 =0.983, RMSE=0.062, MAE=0.19, meeting the accuracy requirements for engineering applications (Kt error <5%). The training set consists of 2400 groups and the test set of 600 groups, randomly divided in an 8:2 ratio.
[0067] Finite element simulations were performed using Abaqus 2022 software. The material constitutive model was an isotropic linear elastic model. A uniaxial tensile load was applied up to 50% of the yield strength. The stress concentration factor was defined as the ratio of the maximum principal stress in the defect neighborhood to the nominal tensile stress far from the defect influence zone. .
[0068] Furthermore, the step of constructing a comprehensive quantitative score based on equivalent volume and stress concentration factor, and outputting the defect severity level according to a preset grading threshold, includes: performing a product operation on the normalized equivalent volume and stress concentration factor to obtain the basic damage base; performing an operation on the normalized stress concentration factor combined with the material sensitization coefficient to obtain the stress amplification coefficient; performing a linear weighted summation on the basic damage base and the stress amplification coefficient to output the comprehensive quantitative score; and comparing the comprehensive quantitative score with the preset grading threshold within an interval to determine and output the defect severity level.
[0069] Specifically, normalization is performed on the equivalent volume Vdef and the stress concentration factor. Max-min normalization is performed separately: the equivalent volume after normalization. =( - Vmin) / (Vmax - Vmin), where the normalization range is determined by historical sample statistics: Vmin=0mm 3 (Defect-free baseline), Vmax=5mm 3 (The upper limit of the maximum acceptable defect volume is determined by downstream process review); Normalized value of stress concentration factor = ( -1.0) / (4.5 -1.0), normalized range [0,1] corresponds to ∈[1.0, 4.5] (truncated to 1 if it exceeds the upper limit).
[0070] Calculation of the base damage threshold Dbase: Dbase = × This product reflects the comprehensive damage basis of the defect's geometric quantity (volume) and mechanical hazard (stress concentration degree), and its value ranges from [0,1].
[0071] Calculation of stress amplification factor Astress: Astress = [exp(β × )- 1] / [exp(β)-1], where exp is the natural exponential function and β is the material sensitization coefficient, reflecting the sensitivity of the alloy to stress concentration. The β value is determined by material testing: for 6061-T6 aluminum alloy, β=1.5 (at this time, Kt=0.6 corresponds to Astress≈2.46); for 6063-T5 aluminum alloy, β=1.2 (lower strength materials are relatively less sensitive to notches, Kt=0.6 corresponds to Astress≈2.05); for 7075-T6 aluminum alloy, β=1.8 (high strength aluminum alloys are more sensitive to notches, Kt=0.6 corresponds to Astress≈2.94). Method for determining β value: Through fatigue tests of aluminum alloys with standard notches (radius 0.1mm to 2.0mm, a total of 5 types), the linear relationship between the logarithm of fatigue life and stress concentration factor (i.e., the linear regression slope of lg(N)~Kt) is fitted, and the β value is calculated. Each β value is supported by experimental data from at least 30 fatigue specimens.
[0072] The comprehensive quantitative score Scomp is calculated as follows: Scomp = w1×Dbase + w2×Astress, where the linear weighting coefficients w1=0.4 and w2=0.6. The weight values are determined based on the following: Analysis of the fracture surfaces of 100 historically failed aluminum bars was conducted to statistically determine the proportion of failure mechanisms dominated by stress concentration (fracture surface exhibiting a fan-shaped radial crack morphology) versus those dominated by volumetric damage (fracture surface exhibiting a cup-cone ductile fracture morphology). The statistical results show that 62 aluminum bars (58%–64%, mean 62%) failed primarily due to stress concentration, and 38 aluminum bars (36%–42%, mean 38%) failed primarily due to volumetric damage. After rounding, w2=0.6 (stress concentration contribution) and w1=0.4 (volume damage contribution) were set. This proportion was confirmed as valid through correlation analysis (Pearson correlation coefficient r=0.91) between the Scomp calculation results of 50 validation set aluminum bars and the actual failure severity.
[0073] Graded Output: The preset graded thresholds are jointly determined by the product quality standards (referring to GB / T 31975-2015 Aluminum and Aluminum Alloy Extruded Bars) and downstream process requirements, and are divided into four levels: Level 1 (Excellent): Scomp ∈ [0, 0.25) – extremely minor defects, meets all uses, direct release; Level 2 (Good): Scomp ∈ [0.25, 0.50) – minor defects, meets general structural component uses, proceeds to downstream secondary finishing process; Level 3 (Poor): moderate defects, restricted use, downgraded or re-inspected after grinding; Level 4 (Scrap): severe defects, directly scrapped, triggers process traceability alarm. The graded thresholds {0.25, 0.50, 0.75} are determined by the correlation regression analysis between the fatigue test life and Scomp of 100 sample aluminum bars, corresponding to the Scomp value at the fatigue life inflection point, ensuring that the graded thresholds have clear mechanical and physical basis.
[0074] Example 2: A method for visual inspection and classification of surface defects in aluminum rods, comprising: The camera and projection parameters are determined based on the diameter of the aluminum rod. Multi-frequency sinusoidal structured light is projected and three-step phase-shift acquisition is performed to obtain the original image sequence containing deformation stripes. Vibration displacement signals are obtained using a three-axis Doppler vibration measurement device, and sub-pixel-level vibration motion compensation is performed on the original image sequence. Two-dimensional grayscale images and wrapping phase images of the compensated image sequence are extracted, and absolute phase reconstruction is completed by multi-frequency difference phase expansion; based on the structural cursor positioning parameters, physical size calculation is performed on the absolute phase in the cylindrical coordinate system of the aluminum rod to generate a three-dimensional height map of the defect area; A dual-stream Transformer network is used to encode the two-dimensional grayscale image and the three-dimensional height map through intensity stream and depth stream respectively, and outputs a joint feature vector through cross-modal cross-attention fusion. The joint feature vector is input into the aluminum rod defect knowledge graph to perform graph neural network inference, and the visual classification result and confidence score are output. When the confidence score is lower than the preset threshold, eddy current and ultrasonic detection are triggered. The multimodal anomaly features and visual classification results are spatially registered, temporally aligned and aggregated and fused to output the final defect type label. The equivalent volume, depth-to-diameter ratio, and minimum radius of curvature of the defect perimeter are calculated based on the three-dimensional height map. The stress concentration factor is calculated by combining the material compensation coefficient. A comprehensive quantitative score is constructed based on the equivalent volume and stress concentration factor, and the severity level of the defect is output according to the preset grading threshold.
[0075] Further, the step of determining camera and projection parameters based on the aluminum rod diameter, projecting multi-frequency sinusoidal structured light and performing three-step phase-shift acquisition to obtain the original image sequence containing deformation stripes includes: obtaining the diameter parameters of the aluminum rod to be inspected, and calculating the spatial chord length and maximum chord height of the corresponding arc surface of the aluminum rod in combination with a preset single detection center angle; determining the required field of view size based on the spatial chord length, and determining the required detection depth of field based on the maximum chord height, so as to solve for the camera and projection parameters, which include the camera working distance and the projection focal length; generating sinusoidal structured light patterns of multiple spatial frequencies through a digital micromirror device, and performing three-step phase-shift projection on the surface of the aluminum rod at each spatial frequency, synchronously triggering the camera to acquire the reflected image modulated by the surface deformation, and obtaining the original image sequence containing deformation stripes.
[0076] Furthermore, the subpixel-level vibration motion compensation includes: acquiring the three-dimensional real-time vibration displacement signal of the aluminum rod through a three-axis Doppler vibration measurement device, projecting it onto the imaging plane in combination with the camera working distance and projection focal length in the camera and projection parameters, and obtaining the inter-frame global spatial offset; using the first frame image as the position reference, performing pixel grayscale resampling mapping on each frame of the image sequence to be compensated based on the subpixel interpolation model, and outputting the motion-compensated image sequence.
[0077] Furthermore, the multi-frequency difference phase unfolding to complete absolute phase reconstruction includes: performing three-step phase shift projection and acquisition on each of the three multi-frequency sinusoidal structured lights with spatial frequency decreasing relationship, and obtaining the wrapping phase map corresponding to each spatial frequency respectively; performing pixel-by-pixel difference frequency processing on the wrapping phase maps corresponding to adjacent spatial frequencies to generate a fundamental frequency phase map that covers the global field of view and has no phase folding; using the fundamental frequency phase map as the unfolding reference, the order of each frequency stripe is deduced step by step according to the frequency multiple relationship, and pixel-by-pixel stitching is performed in combination with the wrapping phase map of the highest spatial frequency to obtain a continuous absolute phase map.
[0078] Furthermore, the step of performing physical dimension calculation on the absolute phase in the cylindrical coordinate system of the aluminum rod based on the structural beam positioning parameters to generate a three-dimensional height map of the defect area includes: obtaining the structural beam positioning parameters, which include the camera intrinsic parameter matrix, the projection device intrinsic parameter matrix, and the relative pose extrinsic parameters between the camera and the projection device; establishing a structured light triangulation geometric mapping model based on the camera and projection parameters and the absolute phase reconstruction results, converting the phase information into spatial three-dimensional coordinates; extracting the absolute phase difference data and combining it with the reference phase constraint, mapping it to a cylindrical coordinate system with the aluminum rod axis as the central axis, and finally generating a three-dimensional height map of the defect area through coordinate transformation.
[0079] Furthermore, the joint feature vector output through cross-modal cross-attention fusion includes: converting the two-dimensional grayscale image feature matrix output by intensity stream coding and the three-dimensional height map feature matrix output by depth stream coding into query, key, and value matrices through linear mapping, and performing bidirectional cross-modal cross-attention operation: using the two-dimensional grayscale image features as queries and the three-dimensional height map features as keys and values to obtain a depth-guided intensity attention representation; simultaneously using the three-dimensional height map features as queries and the two-dimensional grayscale image features as keys and values to obtain an intensity-guided depth attention representation; concatenating the two in the channel dimension and then performing nonlinear mapping through a multilayer perceptron to output a joint feature vector.
[0080] Furthermore, the execution of graph neural network inference and the output of visual classification results and confidence scores include: mapping the joint feature vector as the initial embedding vector of the defect node to be tested to the aluminum rod defect knowledge graph, the knowledge graph containing defect instance nodes, defect type nodes, process cause nodes and their associated edges; after aggregating the features of nearest neighbor nodes through a graph attention network, calculating the probability distribution of each defect category through a fully connected classification layer Softmax, taking the category with the highest probability as the visual classification result, and taking the highest probability value as the confidence score.
[0081] Furthermore, the spatial registration, temporal alignment, and aggregation fusion include: in the spatial registration dimension, using the global coordinate system of the visual acquisition device as a reference, the two-dimensional scanning coordinates of the eddy current probe and the depth coordinates of the ultrasonic probe obtained by eddy current and ultrasonic detection are uniformly mapped to the global coordinate system through rigid body transformation to obtain the spatial position registration results of multimodal anomaly features; in the temporal alignment dimension, using the visual acquisition timestamp as a reference, the pulse counting information on the aluminum rod conveying device is extracted, the acquisition delay of each mode is calculated, and compensation alignment is performed; in the aggregation fusion dimension, the probability distribution of the visual classification results, the eddy current anomaly impedance amplitude characteristics corresponding to eddy current and ultrasonic detection, and the ultrasonic echo attenuation rate characteristics are respectively converted into basic probability allocation functions, and multi-source confidence is aggregated through the orthogonal sum rule of DS evidence theory, and the proposition with the highest probability is taken as the final defect type label.
[0082] Furthermore, the calculation of the equivalent volume, depth-to-diameter ratio, and minimum radius of curvature of the defect perimeter based on the three-dimensional height map, combined with the material compensation coefficient to measure the stress concentration factor, includes: taking the surface height of the undeformed aluminum rod as zero reference, performing a two-dimensional spatial double integral on the absolute value of the height difference in the defect area to obtain the equivalent volume; extracting the maximum depth value of the defect area, converting the defect projection area into the equivalent circle diameter, and calculating the ratio of the maximum depth value to the equivalent circle diameter to obtain the depth-to-diameter ratio; extracting the three-dimensional surface point cloud data of the defect contour, calculating the principal curvature of each discrete point, and taking the reciprocal of the absolute value of the global maximum principal curvature as the minimum radius of curvature of the defect perimeter; based on material mechanics constraints, constructing a three-dimensional input feature vector with the depth-to-diameter ratio, minimum radius of curvature, and material compensation coefficient, inputting it into a pre-calibrated stress concentration nonlinear regression model to perform mapping, and obtaining the stress concentration factor.
[0083] Furthermore, the step of constructing a comprehensive quantitative score based on equivalent volume and stress concentration factor, and outputting the defect severity level according to a preset grading threshold, includes: performing a product operation on the normalized equivalent volume and stress concentration factor to obtain the basic damage base; performing an operation on the normalized stress concentration factor combined with the material sensitization coefficient to obtain the stress amplification coefficient; performing a linear weighted summation on the basic damage base and the stress amplification coefficient to output the comprehensive quantitative score; and comparing the comprehensive quantitative score with the preset grading threshold within an interval to determine and output the defect severity level.
[0084] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for visual inspection and classification of surface defects in aluminum rods, characterized in that, include: The camera and projection parameters are determined based on the diameter of the aluminum rod. Multi-frequency sinusoidal structured light is projected and three-step phase-shift acquisition is performed to obtain the original image sequence containing deformation stripes. Vibration displacement signals are obtained using a three-axis Doppler vibration measurement device, and sub-pixel-level vibration motion compensation is performed on the original image sequence. Two-dimensional grayscale images and wrapping phase images of the compensated image sequence are extracted, and absolute phase reconstruction is completed by multi-frequency difference phase expansion; based on the structural cursor positioning parameters, physical size calculation is performed on the absolute phase in the cylindrical coordinate system of the aluminum rod to generate a three-dimensional height map of the defect area; A dual-stream Transformer network is used to encode the two-dimensional grayscale image and the three-dimensional height map through intensity stream and depth stream respectively, and outputs a joint feature vector through cross-modal cross-attention fusion. The joint feature vector is input into the aluminum rod defect knowledge graph to perform graph neural network inference, and the visual classification result and confidence score are output. When the confidence score is lower than the preset threshold, eddy current and ultrasonic detection are triggered. The multimodal anomaly features and visual classification results are spatially registered, temporally aligned and aggregated and fused to output the final defect type label. The equivalent volume, depth-to-diameter ratio, and minimum radius of curvature of the defect perimeter are calculated based on the three-dimensional height map. The stress concentration factor is calculated by combining the material compensation coefficient. A comprehensive quantitative score is constructed based on the equivalent volume and stress concentration factor, and the severity level of the defect is output according to the preset grading threshold.
2. The method for visual inspection and classification of surface defects in aluminum rods according to claim 1, characterized in that: The process of determining camera and projection parameters based on the aluminum rod diameter, projecting multi-frequency sinusoidal structured light, and performing three-step phase-shift acquisition to obtain an original image sequence containing deformation stripes includes: obtaining the diameter parameters of the aluminum rod to be inspected; calculating the spatial chord length and maximum chord height of the corresponding arc surface of the aluminum rod in combination with a preset single detection center angle; determining the required field of view size based on the spatial chord length and the required detection depth of field based on the maximum chord height to solve for the camera and projection parameters, which include the camera working distance and the projection focal length; generating a sinusoidal structured light pattern with multiple spatial frequencies using a digital micromirror device, and performing three-step phase-shift projection onto the surface of the aluminum rod at each spatial frequency, simultaneously triggering the camera to acquire the reflected image modulated by the surface deformation, and obtaining an original image sequence containing deformation stripes.
3. The method for visual inspection and classification of surface defects in aluminum rods according to claim 2, characterized in that: The subpixel-level vibration motion compensation includes: acquiring the three-dimensional real-time vibration displacement signal of the aluminum rod through a three-axis Doppler vibration measurement device, projecting it onto the imaging plane in combination with the camera working distance and projection focal length in the camera and projection parameters, and obtaining the global spatial offset between frames; using the first frame image as the position reference, performing pixel grayscale resampling mapping on each frame of the image sequence to be compensated based on the subpixel interpolation model, and outputting the motion-compensated image sequence.
4. The method for visual inspection and classification of surface defects in aluminum rods according to claim 1, characterized in that: The multi-frequency difference phase unfolding to complete absolute phase reconstruction includes: performing three-step phase shift projection and acquisition on each of the three multi-frequency sinusoidal structured lights with spatial frequency decreasing relationship, and obtaining the wrapping phase map corresponding to each spatial frequency respectively; performing pixel-by-pixel difference frequency processing on the wrapping phase maps corresponding to adjacent spatial frequencies to generate a fundamental frequency phase map that covers the global field of view and has no phase folding; using the fundamental frequency phase map as the unfolding reference, the order of each frequency stripe is deduced step by step according to the frequency multiple relationship, and pixel-by-pixel stitching is performed in combination with the wrapping phase map of the highest spatial frequency to obtain a continuous absolute phase map.
5. The method for visual inspection and classification of surface defects in aluminum rods according to claim 1, characterized in that: The method of generating a 3D height map of the defect region by performing physical dimension calculation on the absolute phase in the cylindrical coordinate system of the aluminum rod based on the structural light calibration parameters includes: obtaining the structural light calibration parameters, which include the camera intrinsic parameter matrix, the projection device intrinsic parameter matrix, and the relative pose extrinsic parameters between the camera and the projection device; establishing a structured light triangulation geometric mapping model based on the camera and projection parameters and the absolute phase reconstruction results, and converting the phase information into spatial 3D coordinates; extracting the absolute phase difference data and combining it with the reference phase constraint, mapping it to a cylindrical coordinate system with the aluminum rod axis as the central axis, and finally generating a 3D height map of the defect region through coordinate transformation.
6. The method for visual inspection and classification of surface defects in aluminum rods according to claim 1, characterized in that: The joint feature vector output through cross-modal cross-attention fusion includes: converting the two-dimensional grayscale image feature matrix output by intensity stream coding and the three-dimensional height map feature matrix output by depth stream coding into query, key, and value matrices through linear mapping, and performing bidirectional cross-modal cross-attention operation: using the two-dimensional grayscale image features as queries and the three-dimensional height map features as keys and values to obtain a depth-guided intensity attention representation; simultaneously using the three-dimensional height map features as queries and the two-dimensional grayscale image features as keys and values to obtain an intensity-guided depth attention representation; concatenating the two in the channel dimension and then performing nonlinear mapping through a multilayer perceptron to output a joint feature vector.
7. The method for visual inspection and classification of surface defects in aluminum rods according to claim 1, characterized in that: The execution of the graph neural network inference, outputting visual classification results and confidence scores, includes: mapping the joint feature vector as the initial embedding vector of the defect node to be tested to the aluminum rod defect knowledge graph, which contains defect instance nodes, defect type nodes, process cause nodes and their associated edges; after aggregating the features of nearest neighbor nodes through a graph attention network, calculating the probability distribution of each defect category through a fully connected classification layer Softmax, taking the category with the highest probability as the visual classification result, and using the highest probability value as the confidence score.
8. The method for visual inspection and classification of surface defects in aluminum rods according to claim 1, characterized in that: The spatial registration, temporal alignment, and aggregation fusion include: In the spatial registration dimension, using the global coordinate system of the visual acquisition device as a reference, the two-dimensional scanning coordinates of the eddy current probe and the depth coordinates of the ultrasonic probe obtained by eddy current and ultrasonic detection are uniformly mapped to the global coordinate system through rigid body transformation to obtain the spatial location registration results of multimodal anomaly features; In the temporal alignment dimension, using the visual acquisition timestamp as a reference, the pulse count information on the aluminum rod conveying device is extracted, the acquisition delay of each mode is calculated, and compensation alignment is performed; In the aggregation fusion dimension, the probability distribution of the visual classification results, the eddy current anomaly impedance amplitude characteristics corresponding to eddy current and ultrasonic detection, and the ultrasonic echo attenuation rate characteristics are respectively converted into basic probability allocation functions, and multi-source confidence is aggregated through the orthogonal sum rule of DS evidence theory, and the proposition with the highest probability is taken as the final defect type label.
9. The method for visual inspection and classification of surface defects in aluminum rods according to claim 1, characterized in that: The calculation of the equivalent volume, depth-to-diameter ratio, and minimum radius of curvature of the defect perimeter based on a 3D height map, combined with the calculation of the stress concentration factor using a material compensation coefficient, includes: taking the surface height of the undeformed aluminum rod as zero reference, performing a two-dimensional double integral on the absolute value of the height difference in the defect area to obtain the equivalent volume; extracting the maximum depth value of the defect area, converting the defect projection area into the equivalent circle diameter, and calculating the ratio of the maximum depth value to the equivalent circle diameter to obtain the depth-to-diameter ratio; extracting the 3D surface point cloud data of the defect contour, calculating the principal curvature of each discrete point, and taking the reciprocal of the absolute value of the global maximum principal curvature as the minimum radius of curvature of the defect perimeter; based on material mechanics constraints, constructing a 3D input feature vector with the depth-to-diameter ratio, minimum radius of curvature, and material compensation coefficient, inputting it into a pre-calibrated stress concentration nonlinear regression model to perform mapping, and obtaining the stress concentration factor.
10. The method for visual inspection and classification of surface defects in aluminum rods according to claim 1, characterized in that: The process of constructing a comprehensive quantitative score based on equivalent volume and stress concentration factor, and outputting the defect severity level according to a preset grading threshold, includes: performing a product operation on the normalized equivalent volume and stress concentration factor to obtain the basic damage base; performing an operation on the normalized stress concentration factor combined with the material sensitization coefficient to obtain the stress amplification factor; performing a linear weighted summation on the basic damage base and the stress amplification factor to output the comprehensive quantitative score; and comparing the comprehensive quantitative score with the preset grading threshold within an interval to determine and output the defect severity level.