Industrial surface defect detection method based on multi-scale feature fusion

By combining a multimodal sensor array with an adaptive geometric correction model, high-precision real-time detection and dynamic monitoring of industrial surface defects have been achieved, solving the problems of modal data delay and misjudgment in traditional methods and improving the stability and accuracy of detection.

CN120948475APending Publication Date: 2025-11-14XIAN AERONAUTICAL UNIV +1
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202511406691.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Traditional industrial surface defect detection methods suffer from modal data delays due to single sensors or asynchronous acquisition, making accurate matching difficult. Furthermore, the lack of in-depth modeling leads to misjudgment of defects at complex material interfaces and inaccurate detection due to blurred motion images.

Method used

A multimodal sensor array is used to collect data synchronously in real time. Spatial transformation and scale normalization are performed through an adaptive geometric correction model. A physical model is formed by combining a dual-stream decoding network and defects to achieve cross-modal feature alignment and dynamic reweighting. Motion blur is compensated by optical flow information to extract spatiotemporal evolution features. Finally, defect identification and evaluation are performed through a multi-task learning network.

Benefits of technology

It improves the stability and accuracy of detection, enabling simultaneous detection of microscopic defects and macroscopic deformations in complex environments, supporting real-time monitoring and trend prediction, and enhancing the specificity of defect identification and the adaptability of the production system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120948475A_ABST
    Figure CN120948475A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial surface defect detection method based on multi-scale feature fusion, and the method comprises the steps: collecting the multi-source data of a detected surface in real time through a multi-modal sensor array, and forming a structured data set through time-space synchronization and denoising; constructing an adaptive geometric correction model to realize spatial transformation and scale normalization of multi-scale features, and cooperatively realizing cross-modal alignment and preliminary fusion through texture and physical attribute branches of a double-flow decoding network; dynamically reweighting the fusion features based on a defect physical model, strengthening physical mechanism defect characterization and suppressing interference; combining optical flow compensation and three-dimensional convolution to extract spatio-temporal evolution characteristics, and forming dynamic defect characterization; and finally outputting defect type and severity evaluation through the classification model in combination with the process parameter library. Therefore, the adaptability of the method to a complex industrial environment is enhanced, and the detection stability can be maintained under different materials, illumination conditions and dynamic interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an industrial surface defect detection method, and more particularly to an industrial surface defect detection method based on multi-scale feature fusion. Background Technology

[0002] In the field of industrial surface defect detection, traditional detection systems typically employ a single sensor or asynchronous acquisition schemes, resulting in significant time delays between different modal data (such as visible light images and infrared thermal images). For example, in the detection of high-speed rotating components, the difference in sampling frequency between visible light cameras and laser rangefinders makes it difficult to accurately match defect location and morphology data, directly affecting the accuracy of subsequent monitoring.

[0003] Furthermore, existing methods mostly employ simple splicing or weighted averaging to achieve multimodal data fusion, but lack in-depth modeling of the spatial relationships between different modal features. Taking the surface inspection of metal castings as an example, the spatial misalignment of texture features and geometric deformation features can lead to microcrack defects being misjudged as normal surface structures. This fusion method exhibits significant robustness defects at complex material interfaces.

[0004] Meanwhile, existing methods still rely on a traditional framework, namely static image analysis, which has a weak ability to characterize blurred or deformed defects generated during motion. In the context of continuous stamping production lines, image blurring caused by motion can completely mask minute defects, leading to inaccurate detection results.

[0005] Therefore, there is an urgent need for an industrial surface defect detection method based on multi-scale feature fusion to solve the technical problems existing in the above-mentioned prior art. Summary of the Invention

[0006] This invention overcomes the shortcomings of the prior art and provides an industrial surface defect detection method based on multi-scale feature fusion.

[0007] To achieve the above objectives, the technical solution adopted by this invention is: an industrial surface defect detection method based on multi-scale feature fusion, comprising the following steps:

[0008] S1. Multi-source data of the tested surface is collected in real time synchronously through a multi-modal sensor array, and the multi-source data is subjected to spatiotemporal synchronization and noise reduction processing to form a structured dataset.

[0009] S2. Based on the geometric parameters of the measured surface, an adaptive geometric correction model is constructed to perform spatial transformation and scale normalization on the multi-scale features in the dataset; and input into the dual-stream decoding network. Through the synergistic effect of the texture feature extraction branch and the physical attribute extraction branch, spatial alignment and preliminary fusion of cross-modal features are achieved.

[0010] S3. Based on the physical model of defect formation, the preliminary fusion features are dynamically reweighted to strengthen the defect representation that conforms to the physical mechanism and suppress interference signals. The data is then input into the spatiotemporal fusion network, where optical flow information is combined to compensate for motion blur. The evolution features of the defect in the time and space dimensions are extracted through three-dimensional convolution to form a defect representation containing dynamic information.

[0011] S4. Based on defect characterization, identify and locate defect types through a classification model, and output a defect severity assessment by combining the process parameters of the tested surface with a rule library.

[0012] In a preferred embodiment of the present invention, the multimodal sensor array includes a visible light camera, an infrared thermal imager, a laser speckle projector, and a laser rangefinder. The data between the sensors in the multimodal sensor array are synchronized in time through hardware synchronization and spatially registered through time feature point matching. The multi-source data includes the texture, color, heat distribution, geometric deformation, and dynamic time series data of the surface under test.

[0013] In a preferred embodiment of the present invention, in step S2, the process of constructing the adaptive geometric correction model includes:

[0014] S201. Obtain the real-time working distance between the camera and the surface being measured using a laser rangefinder, and obtain the three-dimensional point cloud data of the surface through calibration using a structured light sensor or industrial camera to calculate the local radius of curvature.

[0015] S202. Based on the real-time working distance and local radius of curvature, establish a transformation matrix from the original image coordinate system to the standard physical coordinate system;

[0016] S203. By minimizing the feature difference between the resampled feature map and the standard template, the parameters of the transformation matrix are dynamically adjusted using the gradient descent method to achieve adaptive correction.

[0017] In a preferred embodiment of the present invention, the process of spatial alignment and preliminary fusion of the cross-modal features includes:

[0018] S211. The texture feature extraction branch uses a Gabor filter bank and a direction-adjustable filter to extract the micro-texture features of the surface under test; the physical property extraction branch extracts the physical property features of defects through color invariance transformation and infrared thermogram gradient analysis.

[0019] S212. Construct a cross-modal correlation matrix and calculate the mutual information value between microtexture features and defect physical property features at spatial locations.

[0020] The formula for the correlation matrix is ​​as follows: In the formula, For spatial location The correlation matrix at the location; Microscopic texture features; The physical properties of the defect; The eigenvalue distribution probability; k is the feature channel index;

[0021] S213. Perform multi-scale fusion of micro-texture features and defect physical property features. The fused features are then compressed through a convolutional layer to generate a preliminary fused feature map. The fusion rule is as follows:

[0022] In the formula, All are learnable weight parameters; For Hadama accumulation.

[0023] In a preferred embodiment of the present invention, the defect formation physical model includes a fracture mechanics model, a heat conduction model, and a critical pressure model for coating blistering; a single model in the defect formation physical model is selected to perform element-wise multiplication on the preliminary fusion features, and an interference suppression term is introduced. By minimizing the ability of non-physical properties, a reweighted feature set is output.

[0024] In a preferred embodiment of the present invention, motion compensation is performed on the feature set by bilinear interpolation, optical flow consistency constraints are introduced, the difference between the feature set after motion compensation and the current feature set is minimized, the spatiotemporal defect evolution features are extracted step by step by three-dimensional convolution to generate the spatiotemporal feature vector of the defect, and the spatiotemporal feature vector is time-series modeled by a long short-term memory network to obtain the defect characterization.

[0025] In a preferred embodiment of the present invention, the classification model employs a multi-task learning network, which branches into a type classification head and a localization regression head. The type classification head outputs the probability distribution of defect types through global average pooling and fully connected layers, while the localization regression head generates a heat map of the defect region through a spatial transformation network and outputs the defect coordinates and bounding boxes by combining non-maximum suppression.

[0026] In a preferred embodiment of the present invention, the defect severity assessment is achieved through a defect severity index, calculated using the following formula:

[0027] ;

[0028] In the formula, The defect area is the defect bounding box output by the localization regression head of the classification model, which is the actual physical area obtained by the transformation matrix from the original image coordinate system to the standard physical coordinate system of the adaptive geometric correction model in step S2. To determine the contrast between the defect and the background, the mean difference in grayscale contrast between the defect region and the surrounding background region is calculated based on the texture features and physical property features contained in the defect representation extracted from the defect representation generated in step S3. The process influence coefficient is derived from the process parameter association rule base. The rule base is based on the mapping relationship between historical process data and defect characterization. It uses a Bayesian network or lightweight decision tree, and combines the defect type (such as metal cracks, coating blistering) output by the classification model with the current process parameters (such as welding temperature, coating curing time) to output the corresponding coefficient. , , , To integrate the weighting coefficients, the system is trained and optimized using historical defect detection data and process verification results to ensure that the contribution of each parameter to the severity assessment matches actual industrial needs. The defect dynamic evolution coefficient is extracted from the defect characterization generated in step S3. Based on optical flow compensation and three-dimensional convolution, the spatiotemporal evolution characteristics of the defect (such as crack propagation speed and coating bubble volume change rate) are obtained and quantified after time-series modeling through a long short-term memory network (LSTM) to reflect the dynamic development trend of the defect. The defect type weighting coefficient is determined by the defect type output by the classification model. It is preset according to the process hazard level of different defects to strengthen the proportion of high-risk defects in the assessment. This is a defect severity indicator, and its value is mapped to a multi-level severity. The value is mapped to a multi-level severity of minor / moderate / severe / fatal according to a preset threshold. The mapping rule needs to be associated with the defect-process risk correspondence table in the process parameter association rule library.

[0029] In a preferred embodiment of the present invention, the process parameter association rule base is constructed through the mapping relationship between historical process data and defect characterization, and a Bayesian network or lightweight decision tree model is used to realize the reasoning from defect type and severity to process parameter adjustment suggestions.

[0030] This invention addresses the shortcomings of the prior art and has the following beneficial effects:

[0031] (1) Multi-source data such as texture, heat distribution, and geometric deformation are collected in real time by a multi-modal sensor array, and a structured dataset is formed by combining spatiotemporal synchronization and noise reduction processing. This collaborative working mode enhances the adaptability of the method to complex industrial environments and can maintain detection stability under different materials, lighting conditions and dynamic interference.

[0032] (2) The calibration model constructed based on the geometric parameters of the surface under test can perform spatial transformation and scale normalization on multi-scale features, effectively solving the problem of feature loss caused by scale differences in traditional methods. This mechanism ensures the ability to detect micro-defects and macro-deformations simultaneously.

[0033] (3) The dual-stream decoding network achieves spatial alignment and preliminary fusion of cross-modal features through the synergistic effect of the texture feature branch and the physical attribute branch. Combined with the dynamic reweighting mechanism of the defect physical model, it can strengthen the defect representation that conforms to physical laws, suppress irrelevant interference signals, and improve the specificity of defect recognition.

[0034] (4) The spatiotemporal fusion network compensates for motion blur by using optical flow information and extracts the evolution characteristics of defects in the temporal and spatial dimensions by combining three-dimensional convolution. This technology can form a defect characterization containing dynamic information, support real-time monitoring of defects in motion and prediction of their development trends, and provide data support for preventive maintenance. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0036] Figure 1 This is a flowchart of a preferred embodiment of the present invention. Detailed Implementation

[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0038] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein. Therefore, the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0039] Industrial surface defects are complex and easily affected by environmental interference. Single-modal data (such as visible light images) can easily lead to the missed detection of minute defects due to differences in material reflection or changes in illumination. A multi-modal sensor array composed of visible light, infrared, structured light, and laser rangefinders can simultaneously acquire multi-dimensional information such as texture, thermal distribution, and geometric deformation. Visible light captures detailed textures, infrared detects thermal anomalies, and structured light and rangefinders solve the scale distortion problem caused by varying distances. Hardware-synchronized triggering and feature point matching achieve nanosecond-level temporal alignment and micrometer-level spatial registration, providing high-precision, highly complementary structured data input for subsequent feature fusion, effectively improving the robustness of detecting cross-material composite defects.

[0040] Specifically, such as Figure 1 As shown, the industrial surface defect detection method based on multi-scale feature fusion includes the following steps:

[0041] S1. Multi-source data from the tested surface is acquired in real time synchronously using a multi-modal sensor array. This multi-source data undergoes spatiotemporal synchronization and noise reduction processing to form a structured dataset. The multi-modal sensor array includes a visible light camera, an infrared thermal imager, a laser speckle projector, and a laser rangefinder. Data from each sensor in the array is synchronized in time via hardware synchronization and spatially registered through time feature point matching. The multi-source data includes the texture, color, thermal distribution, geometric deformation, and dynamic temporal data of the tested surface.

[0042] Because changes in workpiece position or surface deformation during industrial inspection can lead to inconsistent feature map scales, traditional affine transformations are ill-suited to complex geometric structures. Therefore, step S2 is performed to construct an adaptive geometric correction model based on the geometric parameters of the measured surface, performing spatial transformation and scale normalization on the multi-scale features in the dataset. This model is then input into a dual-stream decoding network, where the synergistic effect of the texture feature extraction branch and the physical attribute extraction branch achieves spatial alignment and preliminary fusion of cross-modal features.

[0043] Specifically, by acquiring 3D geometric parameters (such as working distance and local curvature) in real time through laser ranging and structured light, and combining this with deformable convolution to dynamically resample multi-scale features, the blurring problem in edge regions of traditional interpolation methods can be solved, achieving micron-level feature alignment. A gated attention mechanism is introduced across material boundary regions, dynamically adjusting the fusion weights based on material labels to avoid pseudo-defects caused by reflection differences. This stage, through the collaborative design of hardware and algorithms, significantly improves the alignment accuracy of multi-scale features in physical space, providing reliable basic features for subsequent analysis.

[0044] Furthermore, in step S2, the construction process of the adaptive geometric correction model includes:

[0045] S201. Obtain the real-time working distance between the camera and the surface being measured using a laser rangefinder, and obtain the three-dimensional point cloud data of the surface through calibration using a structured light sensor or industrial camera to calculate the local radius of curvature.

[0046] S202. Based on the real-time working distance and local radius of curvature, establish the transformation matrix from the original image coordinate system to the standard physical coordinate system. The mathematical expression of the transformation matrix is: In the formula, This is the rotation / scaling factor. , is the translation amount. s is the scale factor. A transformation matrix T is applied to the multi-scale features respectively, and spatial transformation and scale normalization of the feature maps are achieved through bilinear interpolation or deformable convolution, ensuring the consistency of feature scale under different distances or curvatures.

[0047] S203. By minimizing the feature difference between the resampled feature map and the standard template, the parameters of the transformation matrix are dynamically adjusted using the gradient descent method to achieve adaptive correction.

[0048] Furthermore, the adaptive geometric correction model is implemented through a hardware and algorithm collaboration approach. The laser rangefinder and structured light sensor provide real-time geometric parameters, and the algorithm module dynamically generates a transformation matrix based on the parameters and applies it to feature map resampling.

[0049] Multi-scale features include at least three different resolution levels (e.g., original image, 1 / 2 downsampled, 1 / 4 downsampled). The feature maps of each level are independently geometrically corrected and normalized to ensure that the alignment accuracy of each scale feature in physical space is ≤0.1mm.

[0050] The parameters of the transformation matrix are initialized through offline calibration. The calibration process uses a checkerboard target and the least squares method to fit the initial transformation parameters, and dynamic fine-tuning is performed based on the initial parameters during online detection.

[0051] Furthermore, the process of spatial alignment and preliminary fusion of cross-modal features includes:

[0052] S211. The texture feature extraction branch uses a Gabor filter bank and a directional adjustable filter to extract the microscopic texture features of the tested surface, such as the interlayer structure of carbon fiber and the direction of scratches on metal surfaces. The physical property extraction branch extracts the physical property features of defects through color invariance transformation and infrared thermogram gradient analysis, such as the thermal diffusivity of glass concretions and the spectral distribution of coating color differences.

[0053] S212. Construct a cross-modal correlation matrix and calculate the mutual information value of microtexture features and defect physical property features at spatial locations.

[0054] The formula for the correlation matrix is ​​as follows: In the formula, For spatial location The correlation matrix at that location. It represents microscopic texture features. These are the physical properties of the defects. represents the eigenvalue distribution probability. k is the feature channel index. Based on the mutual information maximization criterion, a spatial transformation network is used to... Perform an affine transformation to make it analogous to... Align in spatial location to ensure that the overlap rate of cross-modal features in the defect area is ≥95%.

[0055] S213. Multi-scale fusion of micro-texture features and defect physical property features is performed. The fused features are then compressed through a convolutional layer to generate a preliminary fused feature map. The fusion rule is as follows:

[0056] In the formula, All of these are learnable weight parameters. For Hadama accumulation.

[0057] It is important to note that for cross-material boundary regions of composite materials (such as carbon fiber + glass), a material recognition branch network is introduced. The fusion weights of the dual-stream network are dynamically adjusted through material labels (such as metal, plastic, and glass) to avoid feature distortion caused by differences in the reflective properties of different material surfaces.

[0058] Adaptive optimization across material boundaries is achieved through a gated attention mechanism, where the material recognition branch outputs a gate signal to control the update of the fusion weights of the two-stream network.

[0059] Because purely data-driven models are susceptible to noise interference (such as the similarity between material texture and defects), they struggle to explain defect formation patterns. Therefore, physical models of defect formation, including fracture mechanics and thermal conduction, are introduced. By explicitly constraining feature weights using key parameters such as stress intensity factor and thermal diffusivity, defect regions conforming to physical mechanisms (such as high-stress areas at crack tips) can be accurately located. For dynamic scenarios, optical flow estimation is used to compensate for motion blur, combined with 3D convolution to extract spatiotemporal evolution features, addressing the distortion problem of traditional 2D methods in moving workpiece detection. This stage, through physical-data collaborative driving, improves the detection rate of minute defects and enables real-time tracking of dynamic defects, providing a theoretical basis for process optimization.

[0060] Specifically, in step S3, a physical model is formed based on the defect. The preliminary fusion features are dynamically reweighted to strengthen the defect representation that conforms to the physical mechanism and suppress interference signals. This model is then input into the spatiotemporal fusion network. Combined with optical flow information to compensate for motion blur, the evolution features of the defect in the temporal and spatial dimensions are extracted through three-dimensional convolution to form a defect representation containing dynamic information.

[0061] The physical models for defect formation include fracture mechanics models, heat conduction models, and critical pressure models for coating blistering.

[0062] The formula for the fracture mechanics model is as follows: In the formula, ρ is the stress. L is the crack length. W is the sample width. This is a geometric correction function.

[0063] The formula for the heat conduction model is: In the formula, is the thermal diffusivity. Q is the heat source term. This is the volumetric heat capacity. It indicates how quickly temperature T changes with time t. It describes the curvature of temperature distribution in space, that is, the temperature diffusion caused by heat conduction, where heat is transferred from high temperature regions to low temperature regions.

[0064] The formula for the critical pressure model of coating blistering is as follows: In the formula, It is surface energy. This represents the pressure difference between the inside and outside. This refers to the adhesion between the substrate and the coating. Where is the bubble radius.

[0065] A single model from the physical model of defect formation is selected to perform element-wise multiplication on the initial fused features, and an interference suppression term is introduced. By minimizing the ability of non-physical properties, a reweighted feature set is output. Motion compensation is performed on the feature set through bilinear interpolation, and an optical flow consistency constraint is introduced to minimize the difference between the motion-compensated feature set and the current feature set. The spatiotemporal defect evolution features are extracted stepwise through three-dimensional convolution to generate the spatiotemporal feature vector of the defect. Finally, the spatiotemporal feature vector is temporally modeled through a long short-term memory network to obtain the defect representation.

[0066] Traditional inspection methods only output defect locations, failing to directly guide production adjustments. Therefore, a multi-task learning network (classification + localization) is employed, associated with a rule base for process parameters. By sharing underlying features to reduce computation, Bayesian networks or lightweight decision trees are used to rapidly infer process adjustment suggestions, achieving closed-loop control of "detection-analysis-adjustment." This stage decouples defect severity assessment from process standards, using unified quantitative indicators (such as defect area and contrast) to achieve cross-material / process evaluation. Combined with a real-time feedback mechanism, this shortens process adjustment response time, effectively reducing defect incidence and enhancing the production system's adaptability.

[0067] Specifically, in step S4, based on defect characterization, the defect type is identified and located through a classification model, and the defect severity assessment is output by combining the process parameters of the tested surface with the rule library.

[0068] The classification model employs a multi-task learning network, branching into a type classification head and a localization regression head. The type classification head outputs the defect type probability distribution through global average pooling and fully connected layers, while the localization regression head generates a defect region heatmap through a spatial transformation network and combines non-maximum suppression to output defect coordinates and bounding boxes.

[0069] Defect severity assessment is achieved through a defect severity index, calculated using the following formula:

[0070] ;

[0071] In the formula, The defect area is the defect bounding box output by the localization regression head of the classification model, which is the actual physical area obtained by the transformation matrix from the original image coordinate system to the standard physical coordinate system of the adaptive geometric correction model in step S2. To determine the contrast between the defect and the background, the mean difference in grayscale contrast between the defect region and the surrounding background region is calculated based on the texture features and physical property features contained in the defect representation extracted from the defect representation generated in step S3. The process influence coefficient is derived from the process parameter association rule base. The rule base is based on the mapping relationship between historical process data and defect characterization. It uses a Bayesian network or lightweight decision tree, and combines the defect type (such as metal cracks, coating blistering) output by the classification model with the current process parameters (such as welding temperature, coating curing time) to output the corresponding coefficient. , , , To integrate the weighting coefficients, the system is trained and optimized using historical defect detection data and process verification results to ensure that the contribution of each parameter to the severity assessment matches actual industrial needs. The defect dynamic evolution coefficient is extracted from the defect characterization generated in step S3. Based on optical flow compensation and three-dimensional convolution, the spatiotemporal evolution characteristics of the defect (such as crack propagation speed and coating bubble volume change rate) are obtained and quantified after time-series modeling through a long short-term memory network (LSTM) to reflect the dynamic development trend of the defect. The defect type weighting coefficient is determined by the defect type output by the classification model. It is preset according to the process hazard level of different defects to strengthen the proportion of high-risk defects in the assessment. This is a defect severity indicator, and its value is mapped to a multi-level severity. The value is mapped to a multi-level severity of minor / moderate / severe / fatal according to a preset threshold. The mapping rule needs to be associated with the defect-process risk correspondence table in the process parameter association rule library.

[0072] The process parameter association rule base is constructed by mapping historical process data with defect characteristics, and uses Bayesian networks or lightweight decision tree models to achieve reasoning from defect type and severity to process parameter adjustment suggestions.

[0073] In industrial settings, the surfaces being tested are composed of complex materials (such as metals, glass, and composite materials), exhibit diverse defect types (cracks, stones, debonding), and are susceptible to environmental interference (changes in lighting, motion blur). Traditional methods either rely on single-modal data, leading to missed detections, or employ purely data-driven models that lack interpretability. This invention utilizes a multi-modal sensor array to simultaneously acquire multi-source data, including texture, thermal distribution, and geometric deformation. Combined with hardware synchronization and feature point matching, it achieves nanosecond-level temporal alignment and micrometer-level spatial registration, overcoming the detection blind spots caused by isolated data in traditional methods.

[0074] This invention addresses the problem of workpiece position changes or surface deformation by acquiring three-dimensional geometric parameters in real time through laser ranging and structured light, and dynamically adjusting the feature map scale by combining deformable convolution to ensure that the alignment accuracy of multi-scale features in physical space is ≤0.1mm, providing a reliable foundation for cross-modal fusion.

[0075] By introducing defect formation models based on fracture mechanics and thermal conduction, and by explicitly constraining feature weights using key parameters such as stress intensity factor and thermal diffusivity, defect regions that conform to physical laws (such as high-stress areas at crack tips) can be accurately located. At the same time, false defects caused by material reflection differences can be suppressed, thus significantly improving the detection rate of minute defects.

[0076] For scenarios involving moving workpieces, motion blur is corrected through optical flow compensation, and the evolution characteristics of defects in the temporal and spatial dimensions (such as crack propagation speed and bubble volume change) are extracted by combining 3D convolution. LSTM is then used to model the temporal patterns to achieve real-time tracking and prediction of dynamic defects.

[0077] Based on the preferred embodiments of the present invention described above, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. An industrial surface defect detection method based on multi-scale feature fusion, characterized in that, Includes the following steps: S1. Multi-source data of the tested surface is collected in real time synchronously through a multi-modal sensor array, and the multi-source data is subjected to spatiotemporal synchronization and noise reduction processing to form a structured dataset. S2. Based on the geometric parameters of the surface under test, construct an adaptive geometric correction model and perform spatial transformation and scale normalization on the multi-scale features in the dataset. The data is then input into a dual-stream decoding network, where the spatial alignment and initial fusion of cross-modal features are achieved through the synergistic effect of the texture feature extraction branch and the physical attribute extraction branch. S3. Based on the physical model of defect formation, the preliminary fusion features are dynamically reweighted to strengthen the defect representation that conforms to the physical mechanism and suppress interference signals. The data is then input into the spatiotemporal fusion network, where optical flow information is combined to compensate for motion blur. The evolution features of the defect in the time and space dimensions are extracted through three-dimensional convolution to form a defect representation containing dynamic information. S4. Based on defect characterization, identify and locate defect types through a classification model, and output a defect severity assessment by combining the process parameters of the tested surface with a rule library.

2. The industrial surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that: The multimodal sensor array includes a visible light camera, an infrared thermal imager, a laser speckle projector, and a laser rangefinder. Data between the sensors in the multimodal sensor array is synchronized in time through hardware synchronization and spatially registered through time feature point matching. The multi-source data includes the texture, color, heat distribution, geometric deformation, and dynamic time series data of the surface under test.

3. The industrial surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that: In step S2, the construction process of the adaptive geometric correction model includes: S201. Obtain the real-time working distance between the camera and the surface being measured using a laser rangefinder, and obtain the three-dimensional point cloud data of the surface through calibration using a structured light sensor or industrial camera to calculate the local radius of curvature. S202. Based on the real-time working distance and local radius of curvature, establish a transformation matrix from the original image coordinate system to the standard physical coordinate system; S203. By minimizing the feature difference between the resampled feature map and the standard template, the parameters of the transformation matrix are dynamically adjusted using the gradient descent method to achieve adaptive correction.

4. The industrial surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that: The process of spatial alignment and preliminary fusion of the cross-modal features includes: S211. The texture feature extraction branch uses a Gabor filter bank and a direction-adjustable filter to extract the micro-texture features of the surface under test; the physical property extraction branch extracts the physical property features of defects through color invariance transformation and infrared thermogram gradient analysis. S212. Construct a cross-modal correlation matrix and calculate the mutual information value between microtexture features and defect physical property features at spatial locations. The formula for the correlation matrix is ​​as follows: In the formula, Spatial location The correlation matrix at the location; Microscopic texture features; The physical properties of the defect; The eigenvalue distribution probability; k is the feature channel index; S213. Perform multi-scale fusion of micro-texture features and defect physical property features. The fused features are then compressed through a convolutional layer to generate a preliminary fused feature map. The fusion rule is as follows: In the formula, All are learnable weight parameters; For Hadama accumulation.

5. The industrial surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that: The physical model for defect formation includes a fracture mechanics model, a heat conduction model, and a critical pressure model for coating blistering. A single model from the physical model for defect formation is selected to perform element-wise multiplication on the preliminary fusion features, and an interference suppression term is introduced. By minimizing the ability of non-physical properties, a reweighted feature set is output.

6. The industrial surface defect detection method based on multi-scale feature fusion according to claim 5, characterized in that: Motion compensation is performed on the feature set by bilinear interpolation, and optical flow consistency constraints are introduced to minimize the difference between the feature set after motion compensation and the current feature set. The spatiotemporal evolution features of defects are extracted step by step by three-dimensional convolution to generate spatiotemporal feature vectors of defects. The spatiotemporal feature vectors are then modeled temporally by a long short-term memory network to obtain the defect representation.

7. The industrial surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that: The classification model employs a multi-task learning network, which branches into a type classification head and a localization regression head. The type classification head outputs the probability distribution of defect types through global average pooling and fully connected layers, while the localization regression head generates a heat map of the defect region through a spatial transformation network and outputs the defect coordinates and bounding boxes by combining non-maximum suppression.

8. The industrial surface defect detection method based on multi-scale feature fusion according to claim 7, characterized in that: The severity of the defect is assessed using a defect severity index, calculated as follows: ; In the formula, The defect area is equal to the number of pixels. To enhance the contrast between the defect and the background; This is the process influence coefficient; , , , This is the overall weighting coefficient; The defect dynamic evolution coefficient is extracted from the defect characterization. For defect type weighting coefficients; It is an indicator of defect severity, and its value is mapped to a multi-level severity.

9. The industrial surface defect detection method based on multi-scale feature fusion according to claim 1, characterized in that: The process parameter association rule base is constructed by mapping historical process data with defect representations, and uses Bayesian networks or lightweight decision tree models to achieve reasoning from defect type and severity to process parameter adjustment suggestions.

Citation Information

Cited By

  • Welded pipe surface defect detection method

    CN121208012A

  • Steel plate surface defect detection system based on space-time mutual attention and sparse space-time perception attention

    CN121213558A

  • Method for detecting defects of inner surface and outer surface of titanium alloy pipe

    CN121353583A

  • Metal plate surface defect detection method fusing multi-source sensing data

    CN121499764A

  • Rotating wheel defect detection and evaluation method and related device

    CN121659002A