Metal product surface defect detection method and system based on machine vision

By combining composite sensors with multi-task neural networks, the problems of incomplete information acquisition and data integration in the surface inspection of metal products have been solved, achieving high-precision defect detection and automated decision-making, and improving inspection efficiency and consistency of quality control.

CN122023318AInactive Publication Date: 2026-05-12GUANGZHOU HEYOU MOULD & PLASTIC TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGZHOU HEYOU MOULD & PLASTIC TECH CO LTD
Filing Date
2026-01-27
Publication Date
2026-05-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing surface inspection methods for metal products are difficult to collect comprehensive information in complex environments and are difficult to effectively integrate different types of data, resulting in the loss or misinterpretation of defect features. In particular, under strong light illumination, reflective patches can obscure the true defects, making it difficult to distinguish between various defect types.

Method used

A composite sensor is used to simultaneously acquire 3D point cloud and multi-band images. Through multi-task neural network processing, the image data is decomposed into diffuse reflection and specular reflection components, and targeted filtering is performed. Through multimodal data fusion and neural radiation field model reconstruction, a high-resolution 3D defect model is generated, and finally a structured inspection report is generated.

Benefits of technology

It enables accurate collection and effective integration of various information from metal surfaces in complex environments, significantly improving the accuracy and efficiency of defect detection. It can clearly identify subtle defects and generate structured reports that support production line quality control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122023318A_ABST
    Figure CN122023318A_ABST
Patent Text Reader

Abstract

The invention discloses a metal product surface defect detection method and system based on machine vision, and the method comprises the steps: synchronously collecting three-dimensional geometric and two-dimensional texture information through structured light and a multispectral composite sensor, and obtaining clear initial defect data through a reflection separation technology based on the degree of polarization; precise positioning of sub-pixel-level defect boundaries is realized by using multi-modal data alignment and a contour enhancement technology based on an attention mechanism, multi-task intelligent classification is further performed by fusing contour, geometric and spectral features, and fine distinguishing and quantitative description of defects are completed. A high-fidelity three-dimensional digital model of key defects is constructed through a neural radiation field technology, model self-optimization is realized in combination with view difference and generative completion, multi-dimensional quantitative indexes and a process knowledge rule base are subjected to automatic matching reasoning, a structured detection report is generated, full-process intelligence from perception, analysis and reconstruction to decision making is realized, and the detection efficiency is improved. And the accuracy, efficiency and automation level of industrial quality control are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial quality inspection and intelligent manufacturing technology, specifically a method and system for detecting surface defects in metal products based on machine vision. Background Technology

[0002] Surface defect detection of metal products is an indispensable part of quality control in the manufacturing industry. Its importance lies in its direct relationship with product safety and service life, especially in high-requirement fields such as automobiles and aerospace, where even minor defects can lead to serious consequences. Research in this field is crucial for improving industrial production efficiency and ensuring product quality. However, many current detection methods often fall short in the face of complex environments. For example, on the white body inspection line before automobile body painting, the strong reflective patches formed on the surface of sheet metal parts under strong lighting can easily mask real scratches or dents. Especially under the interference caused by light reflection on the metal surface, misjudgments are prone to occur. At the same time, the ability to distinguish between multiple defect types is insufficient, making it difficult to adapt to diverse production needs. In-depth analysis reveals that the difficulty in inspecting the surface of metal products mainly stems from two core technical factors. First, the comprehensiveness of the collected information is insufficient. Due to differences in material and shape, metal surfaces exhibit vastly different light reflection characteristics. A single acquisition method is insufficient to capture subtle texture changes and three-dimensional morphological information. For example, on the casting surface of an engine block, it is difficult to accurately capture both cracks under dark oxide scale and local protrusions blurred by reflection using only a regular camera. This deficiency further leads to another key problem: when processing data, it is difficult to effectively integrate information from different sources. For instance, the matching between planar images and depth contours often deviates, resulting in the loss or misinterpretation of defect features.

[0003] Specifically, on actual production lines, inspection equipment may miss minute cracks or scratches because it cannot accurately adapt to the reflective properties of metal surfaces. Furthermore, even when multiple data points are collected, the lack of effective integration methods leads to ambiguous judgments of defect types, such as mistaking scratches for dents, affecting subsequent repair decisions. Therefore, how to accurately collect various information from metal surfaces in complex environments and effectively integrate different types of data to accurately distinguish various defects has become a key issue in improving inspection effectiveness. Solving this problem will directly impact quality control and efficiency improvement in the production process. Summary of the Invention

[0004] The purpose of this invention is to provide a method and system for detecting surface defects in metal products based on machine vision, which realizes a complete automated closed loop from accurate detection to intelligent decision-making.

[0005] The objective of this invention can be achieved through the following technical solutions: This application provides a machine vision-based method for detecting surface defects in metal products, comprising the following steps: S1. Simultaneously acquire three-dimensional point cloud and multi-band images through composite sensors, and obtain preliminary information on the spatial location and suspected category of defects through multi-task neural network processing. S2. Decompose the image data in the initial positioning area into diffuse reflection components and specular reflection components, and perform targeted filtering processing on each component to obtain image data that suppresses specular interference. S3. Align the filtered diffuse reflection component image with the 3D point cloud data in terms of spatial coordinates, and calculate the consistency metric of the features of the two. If the consistency metric is lower than the first set threshold, start the contour enhancement network based on the coordinate attention mechanism and output the enhanced sub-pixel contour data of the defects. S4. Input the enhanced subpixel-level contour data of the defects, the corresponding three-dimensional local point cloud, and the spectral feature vector extracted from the diffuse reflection component image into the defect classification network for processing to obtain the classified defect type label. S5. For the classified defect type labels, use multispectral images from multiple perspectives and corresponding 3D point clouds as training data to construct a local neural radiation field model and generate a high-resolution 3D defect model that integrates spectral reflectance characteristics. S6. Compare the new view generated by the local neural radiation field model with the original view by difference. If the difference value exceeds the second set threshold, trigger the image completion network to complete it and back-optimize and update the local neural radiation field model to obtain complete defect multimodal description data.

[0006] It also includes: S7, extracting multi-dimensional quantitative indicators based on the complete defect multimodal description data, matching them with a predefined rule base, and automatically generating a structured inspection report containing defect codes, severity levels, and maintenance suggestions.

[0007] This application provides a machine vision-based surface defect detection system for metal products, applied to a machine vision-based method for detecting surface defects in metal products, comprising: The multimodal data acquisition and localization module is used to simultaneously acquire 3D point cloud and multiband texture images, and perform preliminary defect localization and classification through a multi-task neural network; The reflection interference suppression and image enhancement module is used to decompose the reflection components based on polarization information and perform targeted filtering to eliminate the high-gloss interference on the metal surface. The multimodal contour refinement and enhancement module is used to obtain defect boundaries with sub-pixel accuracy by verifying the feature consistency between two-dimensional images and three-dimensional point clouds and enhancing the contours. A multi-dimensional intelligent classification and quantization module is used to fuse multi-source features and output structured defect labels containing macro-level categories, micro-level attributes, and size levels through a multi-task classification network. The 3D defect reconstruction module is used to construct a 3D defect model that integrates geometric and spectral reflectance characteristics based on neural radiation field technology. The model self-optimization and data completion module is used to obtain complete multimodal description data of defects through view difference, feature completion and model inverse optimization; The knowledge-driven decision-making and report generation module is used to automatically match defect quantification indicators based on the rule base and generate structured inspection reports.

[0008] The beneficial effects of this invention are as follows: By deploying composite sensors and reflection separation technology based on polarization physics, the core problem of missed and misjudged defects caused by incomplete information acquisition and strong light interference on complex metal surfaces in traditional methods has been solved. It has achieved the simultaneous capture of high-precision three-dimensional morphology and multi-band texture information in a single scan, and can effectively strip away specular highlights. Thus, in typical scenarios such as automobile wheel hubs and stainless steel plates, clear and reliable initial images and location information of defects have been obtained, laying a precise data foundation for the entire detection process. By establishing a collaborative mechanism of strict alignment, consistency verification and intelligent contour enhancement of multimodal data, the key technical bottlenecks of inaccurate fusion of two-dimensional images and three-dimensional geometric data and blurred contours of low-contrast defects have been solved. This ensures the spatial consistency of texture and geometric information and can automatically trigger targeted sub-pixel level contour enhancement, thereby enabling precise characterization and quantification of the boundaries of difficult-to-identify defects such as shallow dents and weak scratches in stamped parts, significantly improving the accuracy of defect geometric characterization. By adopting a classification architecture that integrates deep fusion of multi-source features with high-fidelity reconstruction of neural radiation fields and a self-optimizing feedback model, the system solves the comprehensive problems of coarse defect type differentiation, lack of in-depth analysis tools, and model instability under interference. This enables the system to perform fine classification and attribute quantification of various defects such as scratches, dents, and cracks, and to generate 3D high-fidelity models that support observation and analysis from any perspective for key defects. At the same time, the system has the ability to self-detect and repair missing data, ensuring that the final output defect description has rich dimensions, intuitive visualization, and strong environmental robustness. By integrating an automated reasoning and decision-making engine based on quantitative indicators and process knowledge rule base, the problem of disconnect between detection results and subsequent production and maintenance actions, as well as the inefficiency of relying on human experience for decision-making, has been solved. The multi-dimensional data output by the detection algorithm is automatically transformed into a structured report containing clear defect codes, severity levels, and specific maintenance guidance, realizing a fully automated closed loop from perception and identification to decision execution. This directly empowers the quality control and process optimization of the production line, and significantly improves the consistency and efficiency of quality control. Attached Figure Description

[0009] To better understand and implement this application, the technical solution is described in detail below with reference to the accompanying drawings.

[0010] Figure 1 This is a schematic flowchart of a machine vision-based method for detecting surface defects in metal products, provided in Embodiment 1 of this application. Figure 2 This is a schematic diagram of a machine vision-based surface defect detection system for metal products, provided in Embodiment 2 of this application. Detailed Implementation

[0011] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, exemplary embodiments will be described in detail below, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.

[0012] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0013] The following detailed description of the specific implementation methods, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided in detail.

[0014] Example 1 Please see Figure 1 This embodiment provides a machine vision-based method for detecting surface defects in metal products, including the following steps: S1. The metal surface is scanned by a composite sensor integrating a structured light projection unit and a multispectral imaging unit, and high-precision three-dimensional point cloud data and multi-band surface texture images are acquired simultaneously. The three-dimensional point cloud and multi-band images are processed in parallel by a multi-task convolutional neural network to output the preliminary spatial location region of the defect and the corresponding suspected defect category.

[0015] Step S1 is used to address the problem of insufficient comprehensiveness of the collected information. For example, when inspecting aluminum alloy wheels with complex curved surfaces, a single scan can simultaneously obtain their three-dimensional contour data and surface textures under different wavebands, providing a multi-dimensional information basis for subsequent analysis.

[0016] Further, step S1 specifically includes: A composite sensor scans simultaneously to acquire 3D point clouds generated by structured light projection and multi-band images captured by a multispectral imaging unit. The 3D point cloud is denoised and voxelized to obtain a voxel mesh. The multi-band images are radiometrically corrected and registered to form a standardized multi-channel image. In the online inspection of automotive aluminum alloy wheel hubs, a typical workpiece with complex curved surfaces and bright surfaces, the process includes: using statistical filtering to remove discrete flying point noise caused by dust in the workshop environment and strong specular reflection from the wheel hub itself; then, through voxelized mesh downsampling, while preserving key geometric features such as spokes and rims, converting irregular point clouds into regular voxel meshes, which greatly reduces the computational complexity of subsequent 3D convolution and provides a uniform data structure; and then performing radiometric correction and registration on multi-band images to form standardized multi-channel images. Specifically, for multiple narrow-band images captured by the multispectral imaging unit, a radiometric calibration method based on a standard whiteboard is used to correct the brightness gradient effect caused by uneven lighting in the production line, ensuring the authenticity of texture information; subsequently, through feature point matching and perspective transformation, all band images are strictly registered to the same pixel coordinate system to form a standardized image cube containing multi-dimensional spectral information including visible light and near-infrared light.

[0017] The voxel mesh is input into the point cloud encoding branch, and geometric structure features are extracted through a 3D convolutional neural network; the multi-channel image is input into the image encoding branch, and weighted texture spectral features are extracted through a 2D convolutional neural network combined with a channel attention mechanism. The two features are then fused through a cross-attention module to generate a unified shared feature representation. In the feature extraction stage, geometric structural features are extracted using a 3D convolutional neural network. This involves inputting a voxel grid into the 3D CNN, where the network automatically learns and extracts local and global 3D shape patterns related to defects through multi-layer 3D convolution and pooling operations. Examples include spatial geometric anomalies caused by minor dents, protrusions, or bumps on the wheel hub surface. Simultaneously, weighted texture spectral features are extracted using a 2D convolutional neural network combined with a channel attention mechanism. This involves inputting a standardized multi-channel image into a 2D CNN, where the network not only extracts spatial texture features but also dynamically evaluates the importance of different spectral channels to the current detection task through a channel attention module (such as the SE-Net module). For example, it automatically strengthens bands sensitive to scratches and weakens channels severely affected by ambient light interference, thereby obtaining more robust and discriminative weighted texture-spectral fusion features.

[0018] Based on shared features, a defect probability voxel heatmap is generated by regression through a 3D convolutional decoding head, and the spatial voxel region of the defect is determined by thresholding; at the same time, the initial probability of the region belonging to various types of defects is output by a classification decoding head as the suspected defect category. The process involves generating a defect probability voxel heatmap through regression using a 3D convolutional decoding head. Thresholding is then used to determine the spatial voxel region of the defect. Specifically, the decoder predicts the probability of a defect at each voxel location, generating a 3D heatmap. An empirical threshold is set, and voxels with probabilities higher than this value are marked, thus accurately defining the suspected defect area in the 3D voxel space, such as locating a 3D abnormal block on a wheel hub flange. Simultaneously, a classification decoding head outputs the initial probability of the area belonging to various defect types as the suspected defect category. This refers to another parallel classification branch, based on the same shared features, outputting the probability distribution of the area belonging to predefined categories such as "scratches," "dents," and "stains," providing a preliminary type judgment tendency for subsequent steps. For example, it might determine that the area has a 70% probability of being a scratch.

[0019] Based on the sensor calibration parameters, the spatial voxel region is mapped to the actual workpiece coordinate system to obtain the three-dimensional spatial location of the defect. The three-dimensional spatial location is then associated with the suspected defect category, and a preliminary defect description containing the three-dimensional location and category tendency is output.

[0020] Specifically, by simultaneously acquiring three-dimensional and multispectral data through composite sensors and combining them with multi-task convolutional neural networks for multimodal feature fusion and decoding, the problem of incomplete information from a single sensor is effectively solved. This enables rapid and preliminary three-dimensional localization and coarse screening of defects on complex metal surfaces (such as wheel hubs), providing accurate and reliable spatial location and type tendency basis for subsequent fine analysis.

[0021] S2. Based on the polarization degree information calculated from the multispectral image, the image data in the preliminary positioning area is decomposed into diffuse reflection component and specular reflection component. The specular reflection component is suppressed by nonlocal mean filtering based on regional variance, while the diffuse reflection component is edge-preserving and denoising by anisotropic diffusion filtering, thus obtaining multimodal data after separation of reflection interference.

[0022] Step S2 is used to solve the problem of light reflection interference. For example, on the surface of a bright stainless steel sheet, this method can effectively separate and suppress the specular highlights caused by ambient light and restore the true texture of defects that are covered by highlights.

[0023] Further, step S2 specifically includes: For the initially located defect area, the polarization degree of each pixel is calculated from the corresponding multispectral image. Based on the statistical distribution of polarization degree, the Gaussian mixture model is used to cluster and decompose the pixels in the area into diffuse reflection component and specular reflection component, resulting in two sets of separate image data. For the separated specular reflection component image, the pixel variance of its local region is calculated. Based on the regional variance, the search window and attenuation parameter of the nonlocal mean filter are dynamically adjusted. The high variance (strong specular highlights) region is smoothed and suppressed to obtain the specular suppressed image. Anisotropic diffusion filtering is applied to the separated diffuse reflection component image. The diffusion intensity is controlled according to the image gradient. While smoothing the noise inside the image, the edge information of the defects is protected and sharpened, resulting in a diffuse reflection image with enhanced details. The specular suppression image and the detail-enhanced diffuse reflection image are pixel-weightedly fused, and the fused image is then subjected to adaptive histogram equalization to improve the overall contrast and brightness uniformity, resulting in clear multimodal image data with reflection interference effectively separated and suppressed.

[0024] Furthermore, the degree of polarization (DoLP) of each pixel is calculated from the corresponding multispectral images. This includes: in the inspection of bright stainless steel sheets, firstly, by rotating a polarizer or using a split-focus plane polarization camera, acquiring at least three multispectral images with different polarization directions (e.g., 0°, 45°, 90°) from the same viewpoint. Then, using the Stokes vector formula, the DoLP of each pixel is calculated based on the light intensity values ​​in each direction. This converts the light intensity information into polarization information sensitive to reflection characteristics, providing a crucial basis for subsequently distinguishing between stable diffuse reflection determined by surface microstructure and highly variable specular reflection determined by viewing angle and light source position.

[0025] The Gaussian Mixture Model (GMM) is used to cluster and decompose the pixels within a region into diffuse and specular components. This involves: based on the polarization values ​​of all pixels in the initially located region, their statistical distribution typically exhibits a bimodal characteristic (one peak corresponds to diffuse pixels with low polarization, and the other peak corresponds to specular pixels with high polarization). The GMM is used to fit this bimodal distribution, and the parameters of each Gaussian component are estimated using the expectation-maximization algorithm. Based on this, the posterior probability of each pixel belonging to either the diffuse or specular component is calculated. Finally, each pixel is forcibly assigned to a component based on the maximum posterior probability, thereby decomposing the original image data into two groups at the pixel level: one group mainly contains diffuse images that include the inherent texture of the object, and the other group mainly consists of drastically changing specular highlights.

[0026] For high-variance (strong specular highlights) regions, targeted smoothing and suppression are performed. This includes: firstly, dividing the separated specular reflection component image into several local regions (e.g., 16x16 pixel blocks) and calculating the pixel intensity variance of each region. The variance value directly reflects the intensity and fluctuation of the specular highlights in that region. For high-variance regions, a larger non-local mean filter search window and a smaller attenuation parameter are automatically adopted. This means that the algorithm will search for similar structures over a wider range for weighted averaging and the attenuation of weights will be more gradual, thereby achieving strong and uniform smoothing of strong specular highlights and significantly reducing their intensity and abruptness. For low-variance regions, a smaller search window and standard attenuation parameter are used to avoid unnecessary smoothing and loss of detail.

[0027] Anisotropic diffusion filtering is applied, and the diffusion intensity is controlled according to the image gradient. This includes processing the separated diffuse reflection component image using anisotropic diffusion equations, with the diffusion intensity controlled by a transfer function related to the local image gradient. In flat areas with small image gradients or areas with uniform noise, the transfer function value is large, allowing for strong diffusion to smooth the noise. In areas with large image gradients (such as the edges of scratches and pits), the transfer function value decreases sharply, thereby suppressing or even stopping the diffusion process and effectively protecting the critical defect edge information from being blurred. This process is iterated at multiple scales, ultimately suppressing internal image noise (such as sensor noise) while sharpening the edge contours of defects and making texture details clearer.

[0028] The specular suppression image and the detail-enhanced diffuse image are pixel-weightedly fused. Adaptive histogram equalization is then applied to the fused image. This process involves pixel-level weighted fusion of the specular image (after non-local mean filtering suppression) and the diffuse image (after anisotropic diffusion filtering enhancement), with weights dynamically allocated based on the intensity of the original specular components. In areas where residual highlights are still significant, the specular suppression image is given higher weight to ensure sufficient highlight suppression; in texture-rich areas, the diffuse image is given higher weight to preserve details. The fused image may still have unsatisfactory contrast; therefore, adaptive histogram equalization is further applied. This algorithm divides the image into multiple small regions, calculates and applies histogram equalization independently within each region, and finally eliminates block boundaries through interpolation. This improves the overall image contrast, making hidden defects (such as fine scratches) under suppressed highlights clearly visible, while avoiding local over-brightness or under-brightness issues that may occur with global equalization, achieving a uniform improvement in brightness and detail.

[0029] Specifically, by using polarization physics-based reflection component separation and targeted filtering techniques, the core problem of masking defect textures caused by strong specular highlights on metal surfaces is effectively solved. This achieves intelligent suppression of highlight areas and robust enhancement of the true texture of defects on bright surfaces (such as stainless steel plates), thereby outputting clear, high-contrast, and detailed multimodal image data, laying a reliable foundation for subsequent accurate analysis and classification.

[0030] S3. Align the filtered diffuse reflection component image with the 3D point cloud data in terms of spatial coordinates. Calculate the consistency measure between the aligned 2D image features and the 3D point cloud geometric features in the initial defect location area. If the consistency measure is lower than the first set threshold, start the contour enhancement network based on the coordinate attention mechanism in this area. Take the aligned multimodal data as input and output the enhanced subpixel contour data of the defect.

[0031] Step S3 is used to address the problem of feature loss or misreading caused by data fusion deviation. For example, for shallow depressions on metal stamping parts, the contrast of the two-dimensional image may be very low, but the three-dimensional point cloud has slight undulations. Consistency judgment can trigger the contour enhancement network to obtain more accurate defect boundaries.

[0032] Furthermore, step S3 specifically includes: S31. Using the intrinsic and extrinsic parameters of the composite sensor, the diffuse reflection image and the three-dimensional point cloud are strictly aligned in spatial coordinates to ensure that each image pixel has a corresponding three-dimensional spatial point, forming a pixel-point cloud aligned data pair. S32. Within the determined initial defect area, extract texture gradient features from the aligned image and surface normal or curvature features from the point cloud, respectively, and calculate the cosine similarity of the two feature vectors as the consistency measure of the area. S33. Compare the consistency metric value with the preset first threshold. If it is lower than the threshold, it indicates that there is an inconsistency between the two-dimensional texture and three-dimensional geometric information in the initial defect area, and the outline may be blurred. Trigger the outline enhancement network based on the coordinate attention mechanism. Take the aligned image patch and local point cloud features as input, and focus on the feature conflict area through the attention mechanism. S34. The contour enhancement network outputs a subpixel-level contour probability map. Through nonmaximum suppression and subpixel interpolation techniques, the precise defect boundary pixel coordinates are extracted from the probability map to generate enhanced subpixel-level defect contour data.

[0033] Furthermore, the diffuse reflection image and the 3D point cloud are strictly aligned in spatial coordinates. This includes: in the scenario of inspecting stamped door panels of automobiles, using the pre-calibrated internal parameters (such as camera focal length and distortion coefficient) and external parameters (relative position and attitude between the structured light projector and the multispectral camera) of the composite sensor, an accurate imaging geometric model is established. Through the reprojection method, each point in the 3D point cloud is calculated and mapped to the 2D pixel coordinate system of the diffuse reflection image according to its 3D coordinates and the camera model, ensuring that each 3D point corresponds precisely to a pixel position. For image pixels that do not have a direct 3D point corresponding to them, their 3D coordinates are estimated from the neighboring point cloud data through bilinear interpolation. Finally, a strictly aligned pixel-point cloud data pair with a one-to-one mapping between pixel coordinates and 3D spatial coordinates is formed, laying a spatial foundation for subsequent multimodal feature comparison and fusion.

[0034] Surface normal or curvature features are extracted from the point cloud, and the cosine similarity between the two feature vectors is calculated as a consistency measure for the region. This process includes: In the initially located shallow depression region, firstly, the gradient magnitude and direction of each pixel are calculated from the aligned diffuse image using the Sobel or Prewitt operator to form an image gradient feature vector representing texture abrupt changes; simultaneously, from the corresponding local 3D point cloud, the local tangent plane of each point is fitted based on the PCA (principal component analysis) method, and its normal vector is calculated, or the Gaussian curvature and mean curvature in the neighborhood of the point are further calculated to form a feature vector representing surface geometric deformation. The feature vectors calculated for all pixels / points in the region are averaged or extracted using principal components to obtain a representative image gradient feature vector and a point cloud geometric feature vector. Finally, the cosine similarity between these two vectors is calculated. The closer the value is to 1, the more consistent the texture edge of the image with the geometric deformation of the 3D surface in terms of position and trend. If the value is low (e.g., the image edge is obvious but the point cloud is flat, or the point cloud is undulating but the image texture is smooth), it indicates that there is inconsistency between the two, suggesting that the initially detected contour may be blurry or inaccurate.

[0035] A contour enhancement network based on a coordinate attention mechanism is triggered. Taking aligned image patches and local point cloud features as input, the network focuses on regions with conflicting features. Specifically, when the consistency metric falls below a threshold (indicating that the contour information of a region, such as a shallow depression, is unreliable), the system automatically activates the contour enhancement network. The core of this network is a coordinate attention module, which not only captures cross-image channel information interaction but also establishes long-range dependencies on spatial locations (coordinates). The network inputs aligned image patches (providing texture information) and corresponding point cloud geometric feature maps (providing 3D shape information). The attention mechanism first analyzes the spatial differences between the image feature map and the point cloud feature map, automatically calculating an attention weight map. This weight map highlights regions with conflicting features (i.e., locations with weak edge response but large point cloud curvature changes, which are the potential true boundaries of blurred contours). The network uses these weights to recalibrate and fuse features, guiding subsequent convolutional layers to focus on learning and strengthening the contour features of these key regions, thereby effectively improving the representation ability of low-contrast or blurred defect boundaries.

[0036] By employing nonmaximum suppression and subpixel interpolation techniques, precise defect boundary pixel coordinates are extracted from the probability map. This process includes: the contour enhancement network ultimately outputs a contour probability map with the same size as the input image patch, where each pixel value represents the probability that the point belongs to the defect boundary; nonmaximum suppression is performed on the probability map: for each pixel, the probability values ​​of neighboring pixels are checked along its gradient direction (or the direction of the most drastic probability change). If the current pixel is not a local maximum in that direction, its probability is set to zero. This step effectively suppresses the response of broad boundaries, retaining only probability ridges to obtain candidate boundaries with a single pixel width; subpixel interpolation is then performed: for each retained candidate boundary pixel, the probability value distribution within a small neighborhood (e.g., 3x3) centered on it is used to calculate the extreme points (subpixel precision locations) of the probability distribution through methods such as quadratic surface fitting. The coordinates of these extreme points are the precise boundary locations, with an accuracy of up to one-tenth or even higher than that of a pixel. Finally, this series of subpixel coordinate points are connected sequentially to generate enhanced, high-precision subpixel-level defect contour data, giving clear and accurate digital boundaries to defects that are difficult to distinguish, such as shallow depressions.

[0037] Specifically, by using a multimodal data strict alignment and feature consistency verification mechanism, combined with a contour enhancement network based on coordinate attention, the problem of blurred defect contours caused by the fusion deviation between two-dimensional texture and three-dimensional geometric information is effectively solved. This enables sub-pixel level precise boundary positioning and enhancement of low-contrast defects (such as shallow depressions in stamped parts), significantly improving the accuracy and reliability of defect geometric representation.

[0038] S4. Input the enhanced subpixel-level contour data of the defects, the corresponding three-dimensional local point cloud, and the spectral feature vector extracted from the diffuse reflection component image into the defect classification network. The network includes a shared feature extraction backbone and multiple parallel classification heads. The classification heads are used to output the macroscopic category, microscopic morphological attributes, and size range of the defects, respectively, to obtain the classified defect type label.

[0039] Step S4 aims to improve the ability to distinguish between various defect types and adapt to diverse production needs. For example, in the inspection of automotive parts, it can accurately distinguish scratches (macroscopic category), their roughness (microscopic morphological attribute), and the size class to which their length belongs, providing a precise basis for subsequent process adjustments.

[0040] Furthermore, step S4 specifically includes: S41. For the defect subpixel contour data, calculate the shape descriptor and generate the contour feature vector; for the corresponding three-dimensional local point cloud, calculate its surface geometric statistical features and generate the geometric feature vector; for the corresponding region in the diffuse reflection image, calculate its multi-band spectral statistical features and generate the spectral feature vector; concatenate the contour feature vector, geometric feature vector and spectral feature vector to construct a comprehensive feature vector for characterizing the defect. The calculation of shape descriptors and generation of contour feature vectors includes: utilizing the precise sub-pixel-level contour point sequence of the defect to calculate a set of mathematical descriptors that characterize its overall shape and local features. For example, calculating the Hu moments of the contour (invariant to translation, rotation, and scaling) to describe its macroscopic shape category (whether it is linear, arc-shaped, or complexly branched); calculating the Fourier descriptors of the contour to transform the contour's boundary sequence to the frequency domain, capturing the periodic or complex fluctuation details of its contour; and possibly calculating statistical geometric features such as the contour's compactness and eccentricity. These scalar values ​​are combined into a one-dimensional vector, the contour feature vector, which abstractly represents the two-dimensional planar shape attributes of the defect, helping to distinguish scratches (thin lines) from pits (nearly circular or blocky) in terms of shape.

[0041] Generating geometric feature vectors: Within the 3D point cloud region (i.e., the 3D surface corresponding to the scratch) indexed by the contour data, a series of geometric statistical analyses are performed. The calculations include: the mean, standard deviation, skewness, and kurtosis of the point cloud height, used to describe the overall depth of the scratch, the dispersion of the depth distribution, asymmetry (e.g., one side of the scratch is steeper), and sharpness; calculating the roughness parameters of the local surface (e.g., root mean square height); and also including features based on normal variation, such as the average rate of change of the normal vector, to quantify the steepness or smoothness of the surface. These statistics constitute the geometric feature vector, which quantifies the shape and undulation characteristics of the defect in 3D space.

[0042] The multi-band spectral statistical characteristics are calculated to generate a spectral feature vector: pixel values ​​within the defect region are extracted from a diffuse reflectance multispectral image that is strictly aligned with the contour and point cloud and has suppressed specular reflection. For each spectral band (e.g., the R, G, B channels of visible light and the near-infrared channel), the mean (reflecting the average reflectance intensity), standard deviation (reflecting the uniformity of the internal spectral texture), and ratio or normalization exponent (e.g., (R-NIR) / (R+NIR)) of the pixels within that region are calculated. These cross-band statistical values ​​and relationships are organized into a spectral feature vector that encodes the spectral reflectance characteristics of the material in the defect region, helping to identify spectral anomalies caused by material exposure, oxidation, or contamination due to scratches.

[0043] S42. Input the comprehensive feature vector into a shared deep feature extraction backbone network. The backbone network is composed of fully connected layers or one-dimensional convolutional layers and is used to perform nonlinear transformation and deep abstraction on the input comprehensive feature vector to output a high-dimensional, shared high-level feature representation that integrates multi-source information. The backbone network typically consists of multiple fully connected layers (or one-dimensional convolutional layers) stacked together. Between layers, there are non-linear activation functions (such as ReLU) and optional batch normalization layers. Taking the feature vector of wheel hub scratches as an example, when this vector flows through the backbone network, each layer performs non-linear transformation and feature recombination on it, learning and abstracting deeper and more complex correlation patterns between different modal features. For example, it associates seemingly independent clues such as slender contours, depth changes in specific directions, and spectral anomalies caused by exposed substrates to form the essential high-order representation of scratches. This process removes redundancy and noise from the original features, and finally outputs a high-dimensional, highly fused, and more discriminative shared high-level feature representation, laying the foundation for subsequent fine classification tasks.

[0044] S43. The defect classification network contains three parallel classification heads, which take the shared high-level feature representation as input. The first classification head is a macro-category classifier, which outputs the probability distribution of the defect belonging to each category in the preset macro-category set. The second classification head is a micro-morphological attribute regressor, which outputs continuous numerical attributes used to describe the micro-state of the defect surface. The third classification head is a size interval classifier, which outputs the discrete size level to which the defect belongs in the preset size grading standard. S44. Integrate the output results of the three parallel classification heads, determine the category with the highest probability in the macro category classifier as the final macro category of the defect; use the output value of the micro morphology attribute regressor as the quantitative morphology attribute of the defect; use the output of the size range classifier as the quantitative size range of the defect, and generate a structured defect type label containing multi-dimensional attributes.

[0045] Specifically, by integrating multi-source features such as the precise contours, three-dimensional geometric shapes, and multispectral characteristics of defects, and employing a multi-task classification network for collaborative analysis and decision-making, this method effectively solves the problems of coarse differentiation and singular description of defect types in traditional methods. It achieves refined classification and multi-dimensional quantitative description of various defects such as scratches and pits (e.g., macroscopic categories, microscopic roughness, and size grades), providing accurate and structured data support for subsequent quality assessment and process adjustment.

[0046] S5. For the classified defect type labels, use multispectral images from multiple perspectives and corresponding 3D point clouds as training data to construct a local neural radiation field model, synthesize a new view and corresponding depth map from any perspective within the defect area, and generate a high-resolution 3D defect model that integrates spectral reflectance characteristics.

[0047] Step S5 provides support for in-depth analysis. For example, in the inspection of precision components in aerospace, high-fidelity 3D reconstruction of identified suspected crack defects is performed, which allows engineers to examine its 3D morphology and reflection characteristics from any angle and assists in root cause analysis.

[0048] Furthermore, step S5 specifically includes: S51. Based on the spatial location of the defect that has been classified, extract the multispectral image sub-regions of the defect and the corresponding 3D point cloud subsets from the original multi-view scanning data. Perform spatial registration and scale normalization on the images and point cloud data of all views to construct a multimodal training dataset centered on the defect and with known viewpoints. Among them, a multimodal training dataset centered on defects and with known perspectives is constructed, including: in the periodic inspection of aero-engine turbine blades, once the system classifies suspected crack defects, a high-fidelity reconstruction process is initiated. First, based on the three-dimensional spatial coordinates of the crack on the blade, data from multiple perspectives are extracted from the original detection data. This multi-perspective data may come from two sources: one is that during detection, a composite sensor controlled by a robotic arm or turntable scans the blade from multiple angles; the other is different angle data accumulated from previous detections of the same blade. The system automatically extracts a multispectral image sub-region containing the crack area (e.g., an image patch centered on the crack and extending outwards by a certain number of pixels in visible RGB and near-infrared images) and the corresponding three-dimensional point cloud subset (i.e., the three-dimensional spatial points corresponding to the pixels within the image area). Feature matching (such as SIFT) and iterative nearest-point algorithms are used to spatially register the images and point clouds from all perspectives, unifying them into a common coordinate system centered on the defect, and performing scale normalization to ensure data consistency. Finally, a structured dataset is constructed, where each sample contains: known camera perspective parameters (position and orientation), a cropped multispectral image patch, and a strictly aligned local point cloud. This dataset serves as the fuel for training the Local Neural Radiation Field (NeRF) model.

[0049] S52. Define the boundary cube with the spatial location of the defect, initialize the neural radiation field model in the local space, use the multimodal training dataset as training samples, and use differentiable rendering technology to iteratively train the model parameters with the optimization objective of minimizing the difference between the rendered pixel color and the pixel value of the real multispectral image. At the same time, use the geometric position of the corresponding point cloud as a spatial constraint to make the model implicitly represent the continuous geometric density field and the direction-related spectral reflectance function of the defect region. The optimization objective, achieved through differentiable rendering technology, is to minimize the difference between the rendered pixel color and the pixel value of the real multispectral image. This involves: using the previously constructed dataset as input, initializing a neural radiation field model (essentially a multilayer perceptron) within the local space of the crack (defined as a boundary cube that just encloses the crack). The input is the 3D spatial location and viewing direction, and the output is the volume density and direction-dependent spectral color (RGB or multi-band values) at that location. During training, differentiable rendering technology is used: for each real camera viewpoint in the dataset, light rays are emitted from the camera's optical center to each pixel of the image, passing through the scene. A series of points sampled along the ray path are queried from the neural radiation field model to obtain the density and color. Then, the predicted color of the entire image is synthesized by integrating the classic volume rendering equation. The training objective (loss function) is to minimize the difference between the rendered predicted pixel color and the pixel value of the real multispectral image under all training views (e.g., L2 loss). Through backpropagation, the model parameters are iteratively updated, enabling the model to learn to render an image extremely close to the actual captured image based on the input viewpoint.

[0050] By using the geometric position of the corresponding point cloud as a spatial constraint, the model implicitly represents the continuous geometric density field and direction-dependent spectral reflectance function of the defect region. This includes using the geometric position of the corresponding point cloud as a spatial constraint. A common approach is to add a geometric loss term to the loss function. For example, it forces the model to predict high-density regions (i.e., the object surface) as close as possible to the position of the real 3D point cloud. This guides the model to match the implicit continuous geometric density field (high-density values ​​delineate the object surface) with the measured point cloud data while fitting colors. At the same time, through learning, the model will implicitly construct a direction-dependent spectral reflectance function. That is, it can learn how the crack surface should reflect light of different wavelengths under different viewing angles (light direction), thus reflecting realistic gloss and spectral characteristic changes during rendering.

[0051] S53. After training, the optimized local neural radiation field model is used as a static query function. For any given target camera viewpoint, a high-resolution, realistic multispectral new view of the target camera viewpoint is synthesized by performing volume rendering integration on the continuous field defined by the model. At the same time, based on the cumulative transmittance of light in the density field, an accurate depth map corresponding one-to-one with the pixels of the synthesized view is calculated and generated. Specifically, based on the cumulative transmittance of light in the density field, a precise depth map corresponding to each pixel in the synthesized view is calculated and generated. This includes: during volume rendering, when light passes through the scene, the cumulative transmittance of the light is calculated based on the volume density of each point along the line. The transmittance drops sharply at surfaces (high-density areas). By calculating the distance the light travels from the camera to the point where the transmittance drops to a certain threshold, the surface depth corresponding to that pixel can be obtained. Therefore, while synthesizing a new view, a precise depth map corresponding to each pixel in the synthesized view can be calculated and generated in parallel based on the cumulative transmittance of light in the density field. This depth map provides the precise distance from each pixel to the camera in image form, which is crucial data for subsequent 3D geometric reconstruction and quantitative analysis (such as crack depth and volume).

[0052] S54. Based on the synthesized multi-view depth map, or directly extract the isosurface from the trained neural radiation field density field through the moving cube algorithm, generate a triangular mesh model of the defect surface, associate and map the vertices of the mesh model with the spectral properties predicted by the neural radiation field, and finally output a three-dimensional defect digital model that integrates high-precision geometry and surface spectral reflectance characteristics.

[0053] The process of generating a triangular mesh model for the defect surface includes: utilizing the trained NeRF model itself, since its density field defines the geometry of the space; and employing the moving cube algorithm to extract a fixed density isosurface (usually corresponding to the object surface) from the model's density field to directly extract the triangular mesh model for the defect surface.

[0054] 3D Defect Digital Model: Since the NeRF model essentially learns the spectral emission / reflection intensity of every point and direction in space, the vertices (or faces) of the generated triangular mesh model are mapped back to the NeRF query space. For each vertex, based on its 3D coordinates and normal direction (which can be considered as the viewing direction), the neural radiation field model is queried to predict its spectral properties (such as RGB color or multi-band reflectance values). These properties are then mapped to the vertex color or texture coordinates of the mesh, ultimately outputting a 3D defect digital model that integrates high-precision geometry (from density field constraints and point clouds) and realistic surface spectral reflectance characteristics (from multi-view spectral fitting of NeRF). This model can be imported into 3D analysis software for arbitrary angle measurement, sectioning, lighting simulation, and even used for finite element analysis to evaluate the impact of cracks on structural strength, achieving a leap from 2D detection to 3D digital archiving and in-depth analysis of defects.

[0055] Specifically, by constructing a local neural radiation field model using multi-view spectral and geometric data, the shortcomings of traditional methods in lacking high-fidelity and interactive 3D representation capabilities for key defects are effectively addressed. This enables high-precision implicit characterization of the continuous geometric morphology and direction-related reflection characteristics of classified defects (such as cracks), and can synthesize realistic new views and accurate depth maps from any perspective. Finally, it outputs a 3D model that integrates spectral attributes, providing a powerful tool for in-depth quantitative analysis, visual review, and digital archiving of defects.

[0056] S6. Compare the new view generated by the local neural radiation field model with the original acquired image from the corresponding perspective. If the difference value exceeds the second set threshold in a specific area, it is determined that there is feature loss in that area. Trigger the image completion generative adversarial network based on the difference region and context to complete the lost features and back-optimize and update the local neural radiation field model to obtain complete defect multimodal description data.

[0057] Step S6 is used to ensure the integrity and robustness of the model in complex scenes (such as temporary attachment of oil stains and water droplets), avoid the loss of key features due to temporary occlusion, and ensure the integrity of the final description.

[0058] Furthermore, step S6 specifically includes: Using a local neural radiation field model, for each known original acquisition viewpoint, a corresponding synthetic view is rendered and generated. The pixel-wise absolute difference between the synthetic view and the corresponding original view after processing is calculated in each channel of RGB or multispectral. The process involves calculating the pixel-by-pixel absolute difference between the synthesized view and the corresponding original view after processing. This ensures that the synthesized view, rendered by the neural radiation field model, and the original view, after preprocessing steps (reflection suppression, contour enhancement, etc.), are spatially aligned. A fine-grained comparison is performed on the two images pixel by pixel and channel by channel. For the red, green, and blue channels of the RGB color space, or for each independent band channel of multispectral imaging, the absolute difference of the intensity values ​​at the corresponding pixels is calculated. The process then iterates through each pixel position and each spectral channel in the image, ultimately generating a (or a set of) difference images with the same size and number of channels as the original data. Each pixel value directly quantifies the spectral intensity deviation between the rendered result and the actual observation at that position, providing an accurate data basis for subsequent determination of feature loss areas.

[0059] Connectivity analysis is performed on the difference result image. Connected regions whose average pixel difference value continuously exceeds a pre-set second threshold (determined based on the sensor noise model and reconstruction accuracy requirements) are marked as feature loss regions, and their location coordinates and minimum bounding rectangle range are recorded. Using the original view as a real reference, the marked feature loss region and its surrounding context image within a specified pixel range are used as conditions. The pre-trained image completion model based on generative adversarial networks is then input. Based on contextual semantics and texture information, completion content that is consistent with the original background is generated and filled into the feature loss region to obtain a visually complete and reasonable repaired view. The repaired view is used as a new training sample, and together with the original training data, it is used to perform backpropagation optimization on the local neural radiation field model. The model parameters are iteratively updated to learn the completed features until the average difference between the rendered views of all perspectives and the original / repaired views meets the preset accuracy requirements. All geometric, spectral, and attribute data reconstructed by the optimized model are integrated into a complete defect multimodal description dataset without significant feature loss.

[0060] Specifically, by establishing a feature loss detection mechanism based on the difference between the rendered view and the original view, and combining it with intelligent context completion of generative adversarial networks, the problem of loss of key defect information caused by temporary occlusion such as oil stains and water droplets was effectively solved. This enabled self-verification and self-optimization of the defect reconstruction results of the neural radiation field model, and finally obtained complete, robust and detailed multimodal description data of defects.

[0061] S7. Based on the complete multimodal description data of defects, extract quantitative indicators of multiple dimensions such as defect type, three-dimensional size, depth, contour sharpness and spectral reflectance anomaly. Match the predefined rule base containing material and process knowledge with the quantitative indicators to automatically generate a structured inspection report containing defect code, severity level and maintenance suggestions.

[0062] Step S7 ultimately transforms the technical inspection results into decision-making information that can directly guide production. For example, based on the depth and sharpness of the surface defects of the steel plate, combined with the material process rules, it automatically determines whether the defects are minor scratches that are acceptable or serious scratches that require rework, and provides specific repair areas and method suggestions to achieve a closed loop of quality control.

[0063] Furthermore, step S7 specifically includes: From the complete multimodal description data of defects, five quantitative indicators are extracted and calculated, including: the type of defect, the three-dimensional size representing its space occupation, the maximum depth reflecting its severity, the contour sharpness describing its edge clarity, and the spectral reflectance anomaly representing its material anomaly. The system automatically matches and logically deduces the set of quantitative indicators with a predefined rule base containing material standards and process knowledge; based on the matching results, it automatically outputs a structured inspection report with corresponding defect codes, severity levels, and maintenance suggestions.

[0064] The rule base is defined in the form of "IF (condition) THEN (conclusion)", where the condition part is the threshold or range limitation for the quantitative indicator.

[0065] Specifically, by automatically extracting multi-dimensional quantitative indicators of defects and matching them with a predefined process knowledge rule base, the problem of disconnect between inspection results and production decisions and the inefficiency of relying on human experience judgment is effectively solved. It realizes the automatic transformation from multi-dimensional data to standardized assessment, and finally outputs a structured report containing defect codes, severity levels and specific maintenance suggestions, completing the intelligent closed loop from quality inspection to maintenance guidance.

[0066] Example 2 Please see Figure 2 This embodiment provides a machine vision-based metal product surface defect detection system, applied to a machine vision-based metal product surface defect detection method, including: The multimodal data acquisition and positioning module includes a composite sensor integrating a structured light projection unit and a multispectral imaging unit, used to perform a single scan of the metal surface and simultaneously acquire three-dimensional point cloud data and multi-band surface texture images; it also includes a multi-task convolutional neural network, used to perform parallel processing on the three-dimensional point cloud and multi-band images, and output a preliminary defect description containing three-dimensional spatial location and suspected defect category through feature extraction, fusion and decoding. The reflection interference suppression and image enhancement module is used to process the image data of the initially located defect area. Based on the polarization degree information calculated from the multispectral image, it decomposes the image data into diffuse reflection components and specular reflection components; it then performs nonlocal mean filtering suppression processing based on regional variance on the specular reflection component, and anisotropic diffusion filtering on the diffuse reflection component to perform edge-preserving denoising, finally outputting multimodal image data after reflection interference has been separated and suppressed. The multimodal contour refinement and enhancement module is used to align the filtered diffuse reflection component image and the three-dimensional point cloud data in spatial coordinates, and calculate the consistency metric between the aligned two-dimensional image features and the three-dimensional point cloud geometric features in the initial defect area; when the consistency metric is lower than a set threshold, the contour enhancement network based on the coordinate attention mechanism is activated, and the enhanced sub-pixel contour data of the defect is output with the aligned multimodal data as input. The multi-dimensional intelligent classification and quantization module receives the enhanced defect contour data, the corresponding three-dimensional local point cloud, and the spectral feature vector extracted from the diffuse reflection component image. It includes a shared feature extraction backbone network and multiple parallel classification heads, which are used to output the macroscopic category, microscopic morphological attributes, and size range of the defect, and integrate them to generate a structured defect type label. The 3D defect reconstruction module uses multispectral images from multiple perspectives and corresponding 3D point clouds as training data to construct a local neural radiation field model for the defect region. The model is used to synthesize a new view and corresponding depth map from any perspective within the defect region, and generate a 3D defect digital model that integrates spectral reflectance characteristics. The model self-optimization and data completion module compares the new view generated by the neural radiation field model with the original acquired image by difference. If the difference value exceeds the set threshold, it is determined to be a feature loss region. Then, the image completion model based on generative adversarial network is called to complete the lost features, and the completed data is used to back-optimize and update the local neural radiation field model to obtain complete defect multimodal description data. The knowledge-driven decision-making and report generation module extracts quantitative indicators from multiple dimensions of the complete multimodal defect description data, and matches and logically reasons these quantitative indicators with a predefined rule base containing material and process knowledge to automatically generate a structured inspection report containing defect codes, severity levels, and maintenance suggestions.

[0067] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention are within the scope of the present invention.

Claims

1. A method for detecting surface defects in metal products based on machine vision, characterized in that: Includes the following steps: S1. Simultaneously acquire three-dimensional point cloud and multi-band images through composite sensors, and obtain preliminary information on the spatial location and suspected category of defects through multi-task neural network processing. S2. Decompose the image data in the initial positioning area into diffuse reflection components and specular reflection components, and perform targeted filtering processing on each component to obtain image data that suppresses specular interference. S3. Align the filtered diffuse reflection component image with the 3D point cloud data in terms of spatial coordinates, and calculate the consistency metric of the features of the two. If the consistency metric is lower than the first set threshold, start the contour enhancement network based on the coordinate attention mechanism and output the enhanced sub-pixel contour data of the defects. S4. Input the enhanced subpixel-level contour data of the defects, the corresponding three-dimensional local point cloud, and the spectral feature vector extracted from the diffuse reflection component image into the defect classification network for processing to obtain the classified defect type label. S5. For the classified defect type labels, use multispectral images from multiple perspectives and corresponding 3D point clouds as training data to construct a local neural radiation field model and generate a high-resolution 3D defect model that integrates spectral reflectance characteristics. S6. Compare the new view generated by the local neural radiation field model with the original view by difference. If the difference value exceeds the second set threshold, trigger the image completion network to complete it and back-optimize and update the local neural radiation field model to obtain complete defect multimodal description data.

2. The method for detecting surface defects in metal products based on machine vision according to claim 1, characterized in that: Step S1 specifically includes: A composite sensor scans simultaneously to acquire 3D point cloud and multi-band images; the 3D point cloud is denoised and voxelized to obtain a voxel mesh; the multi-band images are radiometrically corrected and registered to form a standardized multi-channel image. Geometric structure features are extracted from the input point cloud encoding branch of the voxel mesh, and weighted texture spectral features are extracted from the input image encoding branch of the multi-channel image. The two features are then fused through a cross-attention module to generate a shared feature representation. Based on the shared features, a defect probability voxel heatmap is generated by regression through a three-dimensional convolutional decoding head to determine the spatial voxel region of the defect, while a classification decoding head outputs the suspected defect category. Based on the sensor calibration parameters, the spatial voxel region is mapped to the actual workpiece coordinate system, the three-dimensional spatial position is associated with the suspected defect category, and a preliminary defect description is output.

3. The method for detecting surface defects in metal products based on machine vision according to claim 1, characterized in that: Step S2 specifically includes: For the initially located defect area, the polarization degree of each pixel is calculated from the corresponding multispectral image. Based on the statistical distribution of polarization degree, a Gaussian mixture model is used to decompose the region pixels into diffuse reflection components and specular reflection components. For the separated specular reflection component image, the local pixel variance is calculated and the nonlocal mean filter parameters are dynamically adjusted according to the variance to focus on smoothing and suppressing high variance regions. Anisotropic diffusion filtering is applied to the separated diffuse reflection component image, and the diffusion intensity is controlled according to the image gradient to smooth noise and preserve edge information. The two processed images are then pixel-weighted and fused, and the fused image is subjected to adaptive histogram equalization to output clear multimodal image data with suppressed reflection interference.

4. The method for detecting surface defects in metal products based on machine vision according to claim 1, characterized in that: Step S3 specifically includes: By using composite sensor parameters, diffuse reflection images and 3D point clouds are strictly aligned in spatial coordinates to form pixel-point cloud aligned data pairs. Within the initial defect area, texture gradient features are extracted from the aligned image, and surface normal or curvature features are extracted from the point cloud. The cosine similarity between the two feature vectors is calculated as a consistency measure. The consistency metric is compared with a preset threshold. If it is lower than the threshold, a contour enhancement network based on coordinate attention mechanism is triggered. The network takes the aligned image patch and local point cloud features as input and focuses on the feature conflict area through the attention mechanism. The contour enhancement network outputs a subpixel-level contour probability map. It extracts precise defect boundary pixel coordinates through nonmaximum suppression and subpixel interpolation techniques to generate enhanced subpixel-level defect contour data.

5. The method for detecting surface defects in metal products based on machine vision according to claim 1, characterized in that: Step S4 specifically includes: A shape descriptor is calculated to generate a contour feature vector for the defect sub-pixel contour data, a geometric feature vector is generated for the surface geometric statistical features of the three-dimensional local point cloud computing, and a spectral feature vector is generated for the corresponding region of the diffuse reflection image by calculating multi-band spectral statistical features. The three are then concatenated into a comprehensive feature vector. The comprehensive feature vector is input into a shared deep feature extraction backbone network, and a high-level feature representation that integrates multi-source information is output through nonlinear transformation and deep abstraction. The defect classification network contains three parallel classification heads. The first classification head outputs the probability distribution of the macroscopic category of the defect, the second classification head outputs the continuous numerical microscopic morphological attributes, and the third classification head outputs the discrete size level. By integrating the outputs of the three classification heads, the final macroscopic category, quantified morphological attributes, and size range are determined, generating structured multidimensional defect type labels.

6. The method for detecting surface defects in metal products based on machine vision according to claim 1, characterized in that: Step S5 specifically includes: Based on the spatial location of the classified defects, multispectral image sub-regions and 3D point cloud subsets corresponding to the defects are extracted from the original multi-view data. After spatial registration and scale normalization, a multimodal training dataset centered on the defects is constructed. The boundary cube is defined by the spatial location of the defect, the neural radiation field model is initialized, and the training dataset is used to perform iterative training through differentiable rendering technology, with the goal of minimizing the difference between the rendered pixels and the real pixels, while the geometric position of the point cloud is used as a spatial constraint. After training, the optimized model is used as a static query function. By rendering and integrating the continuous field, a high-resolution multispectral new view and the corresponding accurate depth map are synthesized from any target perspective. Based on the synthesized multi-view depth map or by directly extracting the isosurface from the density field, a triangular mesh model of the defect surface is generated, and the vertices are correlated and mapped with the spectral properties predicted by the neural radiation field. Finally, a three-dimensional digital model of the defect that integrates geometric and spectral reflectance properties is output.

7. The method for detecting surface defects in metal products based on machine vision according to claim 1, characterized in that: Step S6 specifically includes: A local neural radiation field model is used to render a synthetic view for each original acquisition viewpoint, and the pixel-wise absolute difference between the synthetic view and the processed original view in each RGB or multispectral channel is calculated. Connectivity analysis is performed on the difference result image. Connected regions whose average pixel difference value continuously exceeds the second threshold are marked as feature loss regions, and their location coordinates and minimum bounding rectangle range are recorded. Using the original view as a real reference, the feature loss region and its surrounding context image are input into a pre-trained generative adversarial completion model. Based on the contextual semantics and texture information, consistent completion content is generated to obtain the repaired view. The repaired view is used as a new training sample, and backpropagation optimization is performed on the local neural radiation field model together with the original data. The model parameters are iteratively updated until the average difference between the rendered view and the original / repaired view meets the accuracy requirements. Finally, a complete defect multimodal description dataset is obtained.

8. The method for detecting surface defects in metal products based on machine vision according to claim 1, characterized in that: S7. Extract multi-dimensional quantitative indicators based on the complete defect multimodal description data, match them with a predefined rule base, and automatically generate a structured inspection report containing defect codes, severity levels, and maintenance suggestions.

9. The method for detecting surface defects in metal products based on machine vision according to claim 8, characterized in that: Step S7 specifically includes: From the complete multimodal description data of defects, five quantitative indicators are extracted and calculated, including: the type of defect, the three-dimensional size representing its space occupation, the maximum depth reflecting its severity, the contour sharpness describing its edge clarity, and the spectral reflectance anomaly representing its material anomaly. The system automatically matches and logically deduces the set of quantitative indicators with a predefined rule base containing material standards and process knowledge; based on the matching results, it automatically outputs a structured inspection report with corresponding defect codes, severity levels, and maintenance suggestions.

10. A machine vision-based surface defect detection system for metal products, applied to the machine vision-based surface defect detection method for metal products as described in any one of claims 1-9, characterized in that: include: The multimodal data acquisition and localization module is used to simultaneously acquire 3D point cloud and multiband texture images, and perform preliminary defect localization and classification through a multi-task neural network; The reflection interference suppression and image enhancement module is used to decompose the reflection components based on polarization information and perform targeted filtering to eliminate the high-gloss interference on the metal surface. The multimodal contour refinement and enhancement module is used to obtain defect boundaries with sub-pixel accuracy by verifying the feature consistency between two-dimensional images and three-dimensional point clouds and enhancing the contours. A multi-dimensional intelligent classification and quantization module is used to fuse multi-source features and output structured defect labels containing macro-level categories, micro-level attributes, and size levels through a multi-task classification network. The 3D defect reconstruction module is used to construct a 3D defect model that integrates geometric and spectral reflectance characteristics based on neural radiation field technology. The model self-optimization and data completion module is used to obtain complete multimodal description data of defects through view difference, feature completion and model inverse optimization; The knowledge-driven decision-making and report generation module is used to automatically match defect quantification indicators based on the rule base and generate structured inspection reports.