Deep learning method for real-time detection of metal additive manufacturing defects

By employing a deep learning approach that integrates multi-source data fusion and multi-algorithm collaboration, real-time detection and process optimization of defects in metal additive manufacturing have been achieved. This addresses the issues of insufficient information dimensions and disconnect between detection and process adjustment in existing technologies, thereby improving the accuracy and adaptability of detection.

CN121744195BActive Publication Date: 2026-07-24GUIYANG VOCATIONAL & TECHNICAL COLLEGE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511896263.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-16
Publication Date
2026-07-24
Estimated Expiration
2045-12-16

AI Technical Summary

Technical Problem

Existing defect detection technologies in metal additive manufacturing suffer from limited information dimensions, lack of interaction between algorithms, disconnect between detection and process adjustment, poor model adaptability, and insufficient real-time performance, making it difficult to meet the online detection needs of industrial production.

Method used

A closed-loop detection system integrating multi-source data fusion and multi-algorithm collaboration is constructed. By simultaneously acquiring infrared temperature, molten pool morphology, acoustic signals, and contact force signals, and employing a deep learning model with improved ResNet50, Transformer, cross-attention mechanism, and feature pyramid network, real-time defect detection and dynamic adjustment of process parameters are achieved.

Benefits of technology

It improves the comprehensiveness of defect feature characterization and the accuracy of detection, realizes real-time active suppression of defects and process optimization, meets the real-time detection needs of industrial production, and reduces labor costs and component scrap rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744195B_ABST
    Figure CN121744195B_ABST
Patent Text Reader

Abstract

The application provides a kind of metal additive manufacturing defect real-time detection deep learning method, comprising: S1. multi-source sensing data acquisition and defect labeling;S2. multi-source data preprocessing and fusion;S3. the construction and training of deep learning model of fusion four algorithms;S4. real-time data acquisition and online preprocessing;S5. defect real-time detection and positioning.The core of the application is to construct a closed-loop detection system of multi-source data fusion and multi-algorithm cooperation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of interdisciplinary technology of defect detection in metal additive manufacturing and deep learning, and in particular to a deep learning method for real-time defect detection in metal additive manufacturing. Background Technology

[0002] Metal additive manufacturing technology forms materials by stacking them layer by layer. During the process, it is easily affected by factors such as fluctuations in laser power, deviations in scanning speed, and changes in powder properties, resulting in defects such as pores, cracks, lack of fusion, and deformation, which seriously affect the mechanical properties and service safety of components.

[0003] Existing defect detection technologies suffer from the following shortcomings: First, they often rely on data from a single sensor, resulting in limited information dimensions and difficulty in comprehensively characterizing defect features. Even when using multi-source data, the fusion is often a simple stitching process, lacking deep interaction between features and dynamic weight adjustment, leading to poor fusion results. Second, deep learning models often employ a single algorithm (such as pure CNN or pure Transformer). CNN struggles to capture global feature correlations, while Transformer lacks sensitivity to local details. Furthermore, the algorithms lack bidirectional interaction, failing to consider both local details and global distribution patterns of defects. Third, detection is disconnected from process adjustments, achieving only "passive identification" of defects without real-time feedback to optimize process parameters, making it difficult to suppress defect expansion or recurrence. Fourth, once trained, the models remain fixed, exhibiting poor adaptability to different components and process scenarios, and detection accuracy is easily affected by environmental changes. In addition, some detection methods suffer from insufficient real-time performance and detection lag, making it difficult to meet the online inspection needs of industrial production. Summary of the Invention

[0004] This invention provides a deep learning method for real-time detection of defects in metal additive manufacturing. The core of this method is to construct a closed-loop detection system that integrates multi-source data fusion and multi-algorithm collaboration.

[0005] To achieve the above objectives, the present invention adopts the following technical solution:

[0006] A deep learning method for real-time defect detection in metal additive manufacturing includes:

[0007] S1. Simultaneously acquire infrared temperature image sequences, molten pool morphology image sequences, acoustic signal sequences, and contact force signal sequences during the metal additive manufacturing process, and perform defect annotation to form an original dataset containing the acquired data and defect annotation information;

[0008] S2. Perform targeted preprocessing on the collected data in the original dataset, perform channel-level fusion of the preprocessed infrared temperature image sequence and the melt pool morphology image sequence, perform dimension-level fusion of the preprocessed acoustic signal sequence and the contact force signal sequence, introduce an attention mechanism to dynamically adjust the weight ratio of the two types of fusion features, generate a unified fusion feature, and associate the fusion feature with the defect annotation information in the original dataset to form a preprocessed dataset.

[0009] S3. Construct a deep learning model that integrates an improved ResNet50, a multi-head self-attention Transformer, a cross-attention mechanism, and a feature pyramid network. The improved ResNet50 is used to extract local spatial features from the fused features. The Transformer is used to extract global correlation features from the fused features. The cross-attention mechanism is used to dynamically fuse local spatial features and global correlation features and output preliminary fused features. The feature pyramid network is used to extract multi-scale features from the preliminary fused features and generate feedback adjustment coefficients through information entropy analysis. The cross-attention mechanism receives the feedback adjustment coefficients, optimizes the fusion weights, and outputs the final fused features. Train the deep learning model based on the preprocessed dataset to obtain the optimal model.

[0010] S4. During the operation of the metal additive manufacturing equipment, real-time infrared temperature image sequence, real-time molten pool morphology image sequence, real-time acoustic signal sequence, and real-time contact force signal sequence are collected simultaneously. The real-time collected data are processed online according to the preprocessing process in step S2, and then real-time fusion features are generated through the fusion method in step S2.

[0011] S5. Input the real-time fused features into the optimal model, and output the real-time core feature map through collaborative reasoning of the four types of algorithms in the optimal model. Based on the real-time core feature map, use the anchor box mechanism to predict the defect location and confidence level. When the confidence level reaches the preset threshold, convert the image coordinates of the defect into the physical coordinates of the equipment forming area to determine that a defect exists.

[0012] The deep learning method for real-time defect detection in metal additive manufacturing, as described in this specification, also includes:

[0013] S6. Defect Classification and Severity Assessment: Based on the real-time core feature map output by the optimal model, the probability distribution of various defects is obtained through fully connected layer mapping to determine the defect type. At the same time, the severity score is output by combining the defect feature intensity and size information. According to the preset threshold, the defects are divided into three levels: mild, moderate and severe. The defect type, physical coordinates, severity level and detection timestamp are integrated to form a detection report.

[0014] The deep learning method for real-time defect detection in metal additive manufacturing, as described in this specification, also includes:

[0015] S7. Feedback Adjustment and Dataset Update: The equipment control system receives the inspection report, executes the corresponding process parameter adjustment strategy according to the defect level, and feeds back the real-time fused features, real-time core feature map, inspection report and adjusted process parameters to the original dataset in step S1 to form the updated dataset. When the number of samples in the updated dataset increases by a preset ratio compared with the original dataset, it returns to step S3 to incrementally train the optimal model to form a closed-loop optimization.

[0016] In this specification, in step S1, infrared temperature image sequence, molten pool morphology image sequence, acoustic signal sequence, and contact force signal sequence are synchronously acquired through an infrared thermal imager, a high-speed camera, an acoustic sensor, and a force sensor during the metal additive manufacturing process. The four types of sensors transmit data through an industrial Ethernet interface using a timestamp synchronization protocol, with the timestamp error controlled within ±1ms. The types of defects marked include pores, cracks, lack of fusion, and deformation. The marking results are linked one-to-one with the sensor data through timestamps.

[0017] In this specification, step S2 includes infrared temperature image preprocessing, which includes 3×3 window midpoint filtering, histogram equalization, and bilinear interpolation scaling; melt pool morphology image preprocessing includes Gaussian filtering with a standard deviation of 1.5, RGB to HSV color space conversion, and bilinear interpolation scaling; acoustic signal preprocessing includes db4 wavelet base 5-layer decomposition wavelet thresholding denoising and fast Fourier transform frequency domain conversion; and contact force signal preprocessing includes 5-window moving average filtering and min-max normalization.

[0018] In this specification, in step S2, channel-level fusion forms a 4-channel image feature, including an infrared temperature image preprocessing result channel, a molten pool morphology image preprocessing result channel, a difference channel and a product channel between the two, and one-dimensional signal data is fused into a one-dimensional feature vector, which is mapped to the same dimension as the flattened 4-channel image feature through a fully connected layer, and then participates in attention-weighted fusion.

[0019] In this specification, in step S3, the improved ResNet50 contains 6 residual blocks. Each residual block consists of two 3×3 convolutional layers connected to one 1×1 shortcut. Each convolutional layer is followed by a batch normalization layer and a ReLU activation function. Finally, a local feature map is output through a max pooling layer with a stride of 2.

[0020] In this specification, in step S3, the cross-attention mechanism first calculates the initial weights based on the global average feature value of the local feature map and the global associated feature map, and then dynamically adjusts the weights by receiving the feedback adjustment coefficients output by the feature pyramid network. The global average feature value is obtained by averaging the feature values ​​at all positions of the feature map.

[0021] In this specification, in step S3, the feature pyramid network generates three feature maps at different scales. After scaling the feature maps at different scales to the same size through bilinear interpolation upsampling, they are spliced ​​together according to the channel dimension to obtain multi-scale fused features. Information entropy analysis is achieved by calculating the probability distribution of feature values ​​of each scale feature map.

[0022] In this manual, in step S5, the anchor frame is set with three scales and three aspect ratios to cover defects of different sizes and shapes. The coordinate transformation is achieved through the calibration matrix of the device coordinate system and the image coordinate system. The calibration matrix is ​​obtained in advance by standard calibration plate.

[0023] In summary, the present invention has at least the following beneficial effects:

[0024] Deep fusion of multi-source data enhances the comprehensiveness of feature representation: By integrating multi-dimensional information from images, acoustics, and force signals through feature-level fusion and attention mechanisms, combined with dynamic weight adjustment, the problem of one-sided information from a single data source is solved, making defect feature representation more accurate.

[0025] Algorithm collaboration enhances detection capabilities: Four types of algorithms interact bidirectionally and work together, capturing local spatial details of defects through CNN, extracting global correlation patterns through Transformer, achieving dynamic fusion of local and global features through Cross-Attention, and then adapting to multi-scale defects through FPN, significantly improving the accuracy and robustness of defect detection.

[0026] A closed-loop optimization system enables proactive defect suppression: a closed-loop process of detection-evaluation-adjustment-optimization is constructed, with detection results directly fed back to the process control system to adjust parameters in a targeted manner to suppress defects. At the same time, the dataset and incrementally trained models are dynamically updated to continuously improve the adaptability of detection and the stability of forming quality.

[0027] Combining real-time performance with practicality: Through edge computing deployment and inference acceleration, it meets the real-time detection needs of industrial production, and can complete the entire process of data collection, defect detection and process adjustment without human intervention, reducing labor costs and component scrap rate. Attached Figure Description

[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 This is a schematic diagram of a deep learning method for real-time detection of defects in metal additive manufacturing involved in this invention.

[0030] Figure 2 This is a schematic diagram of the entire multi-source data processing process involved in this invention.

[0031] Figure 3 This is a schematic diagram of the deep learning model construction and training process involved in this invention.

[0032] Figure 4 This is a schematic diagram of the real-time detection and closed-loop optimization process involved in this invention. Detailed Implementation

[0033] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the embodiments of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.

[0034] The following disclosure provides many different implementations or examples for carrying out different structures of the embodiments of the present invention. To simplify the disclosure of the embodiments of the present invention, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the embodiments of the present invention. Furthermore, reference numerals and / or reference letters may be repeated in different examples of the embodiments of the present invention; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various implementations and / or arrangements discussed.

[0035] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0036] like Figure 1 As shown, this embodiment provides a deep learning method for real-time defect detection in metal additive manufacturing, including the following steps:

[0037] S1. Multi-source sensor data acquisition and defect labeling: Infrared temperature image sequence, molten pool morphology image sequence, acoustic signal sequence, and contact force signal sequence are simultaneously acquired through infrared thermal imager, high-speed camera, acoustic sensor, and force sensor during the metal additive manufacturing process. Defect labeling is completed by combining manual labeling with automatic labeling by industrial CT, forming an original dataset containing the above sensor data and defect labeling information.

[0038] S2. Multi-source data preprocessing and fusion: Targeted preprocessing is performed on various types of sensor data in the original dataset. The preprocessed infrared temperature image sequence and the melt pool morphology image sequence are fused at the channel level, and the preprocessed acoustic signal sequence and the contact force signal sequence are fused at the dimension level. An attention mechanism is introduced to dynamically adjust the weight ratio of the two types of fused features to generate a unified fused feature. This fused feature is associated with the defect annotation information in the original dataset to form a preprocessed dataset.

[0039] S3. Construction and Training of a Deep Learning Model Fuding Four Algorithms: A deep learning model fusing improved ResNet50, multi-head self-attention Transformer, cross-attention mechanism, and feature pyramid network is constructed. The improved ResNet50 is used to extract local spatial features from the fused features, the Transformer is used to extract global correlation features from the fused features, the cross-attention mechanism is used to dynamically fuse local spatial features and global correlation features and output preliminary fused features, the feature pyramid network is used to extract multi-scale features from the preliminary fused features and generate feedback adjustment coefficients through information entropy analysis, the cross-attention mechanism receives the feedback adjustment coefficients to optimize the fusion weights and output the final fused features, and the deep learning model is trained based on the preprocessed dataset to obtain the optimal model.

[0040] S4. Real-time data acquisition and online preprocessing: During the operation of the metal additive manufacturing equipment, real-time infrared temperature image sequence, real-time molten pool morphology image sequence, real-time acoustic signal sequence, and real-time contact force signal sequence are simultaneously acquired by the four types of sensors in step S1. The real-time sensor data is processed online according to the preprocessing process in step S2, and then real-time fusion features are generated through the fusion method in step S2.

[0041] S5. Real-time Defect Detection and Localization: Real-time fused features are input into the optimal model, and the four types of algorithms in the optimal model are used to collaboratively infer and output a real-time core feature map. Based on this real-time core feature map, the anchor box mechanism is used to predict the defect location and confidence level. When the confidence level reaches a preset threshold, the image coordinates of the defect are converted into the physical coordinates of the equipment forming area, and the existence of a defect is determined.

[0042] S6. Defect Classification and Severity Assessment: Based on the real-time core feature map output by the optimal model, the probability distribution of various defects is obtained through fully connected layer mapping to determine the defect type. At the same time, the severity score is output by combining the defect feature intensity and size information. According to the preset threshold, the defects are divided into three levels: mild, moderate and severe. The defect type, physical coordinates, severity level and detection timestamp are integrated to form a detection report.

[0043] S7. Feedback Adjustment and Dataset Update: The equipment control system receives the inspection report, executes the corresponding process parameter adjustment strategy according to the defect level, and feeds back the real-time fused features, real-time core feature map, inspection report and adjusted process parameters to the original dataset in step S1 to form the updated dataset. When the number of samples in the updated dataset increases by a preset ratio compared with the original dataset, it returns to step S3 to incrementally train the optimal model to form a closed-loop optimization.

[0044] In some embodiments, in step S1, the four types of sensors transmit data through an industrial Ethernet interface using a timestamp synchronization protocol. The timestamp error is controlled within ±1ms. The types of defect marking include pores, cracks, lack of fusion, and deformation. The marking results are associated with the sensor data one-to-one through timestamps.

[0045] In some embodiments, in step S2, the infrared temperature image preprocessing includes 3×3 window midpoint filtering, histogram equalization, and bilinear interpolation scaling; the melt pool morphology image preprocessing includes Gaussian filtering with a standard deviation of 1.5, RGB to HSV color space conversion, and bilinear interpolation scaling; the acoustic signal preprocessing includes db4 wavelet basis 5-layer decomposition wavelet thresholding denoising and fast Fourier transform frequency domain conversion; and the contact force signal preprocessing includes 5-window moving average filtering and min-max normalization.

[0046] In some embodiments, in step S2, channel-level fusion forms a 4-channel image feature, including an infrared temperature image preprocessing result channel, a molten pool morphology image preprocessing result channel, a difference channel and a product channel of the two, and one-dimensional signal data is fused into a one-dimensional feature vector, which is mapped to the same dimension as the flattened 4-channel image feature through a fully connected layer, and then participates in attention-weighted fusion.

[0047] In some embodiments, in step S3, the improved ResNet50 contains 6 residual blocks, each residual block consists of 2 3×3 convolutional layers and 1 1×1 shortcut connection, each convolutional layer is followed by a batch normalization layer and a ReLU activation function, and finally a local feature map is output through a max pooling layer with a stride of 2.

[0048] In some embodiments, in step S3, the cross-attention mechanism first calculates the initial weights based on the global average feature value of the local feature map and the globally associated feature map, and then dynamically adjusts the weights by receiving the feedback adjustment coefficients output by the feature pyramid network. The global average feature value is obtained by averaging the feature values ​​at all positions of the feature map.

[0049] In some embodiments, in step S3, the feature pyramid network generates three feature maps at different scales. After scaling the feature maps at different scales to the same size through bilinear interpolation upsampling, they are spliced ​​together according to the channel dimension to obtain multi-scale fused features. Information entropy analysis is implemented based on the probability distribution of feature values ​​of each scale feature map.

[0050] In some embodiments, in step S5, the anchor frame is set with three scales and three aspect ratios to cover defects of different sizes and shapes. The coordinate transformation is achieved through the calibration matrix of the device coordinate system and the image coordinate system. The calibration matrix is ​​obtained in advance by standard calibration plate.

[0051] In some embodiments, in step S7, minor defects do not interrupt equipment operation, only the laser power or scanning speed is adjusted; for moderate defects, equipment operation is paused for 10 seconds and then multi-parameter coordinated adjustment of laser power, scanning speed, and layer thickness is performed; for severe defects, equipment operation is immediately interrupted and an audible and visual alarm is triggered, prompting the operator to check equipment parameters, raw material quality, and equipment status.

[0052] In some embodiments, in step S7, the preset proportion of the number of dataset samples is 10%, the batch size of incremental training is the same as that of the initial training, the number of training rounds is 20% of the number of initial training rounds, the initial learning rate is 50% of the initial training learning rate, and the optimizer uses AdamW and sets the weight decay coefficient to suppress overfitting.

[0053] The technical concept of this invention is as follows:

[0054] First, multi-source data from the manufacturing process is simultaneously acquired using infrared thermal imagers, high-speed cameras, acoustic sensors, and force sensors. After targeted preprocessing, unified fused features are generated using feature-level fusion and attention mechanisms. Then, a deep learning model integrating improved ResNet50, Transformer, Cross-Attention, and FPN is constructed. These four algorithms interact bidirectionally: CNN extracts local features, Transformer captures global correlations, Cross-Attention dynamically adjusts the weights of the two types of features, and FPN extracts multi-scale features and feeds back to optimize weights, forming a collaborative feature extraction capability. After model training and optimization, the model is deployed to edge computing nodes to process multi-source data in real time, completing defect localization, classification, and severity assessment. Finally, based on the assessment results, process adjustment strategies corresponding to mild, moderate, and severe defects are implemented, and real-time data and adjustment parameters are fed back to the dataset, triggering incremental model training to continuously optimize detection performance and achieve real-time, accurate defect detection and active suppression.

[0055] S1. Multi-source sensor data acquisition and defect labeling

[0056] Based on the actual working scenarios of metal additive manufacturing equipment (laser selective melting equipment, electron beam melting equipment), we carry out precise deployment and synchronous data acquisition of multi-source sensors. An infrared thermal imager is fixedly mounted above the observation window of the equipment, with its lens vertically aimed at the molten pool area to ensure complete capture of the molten pool temperature distribution. The sampling frequency is set to 100Hz, and it outputs a grayscale image sequence of 640×480 pixels. The grayscale value of each pixel is linearly correlated with the temperature of the corresponding area. A high-speed camera is deployed on the side of the equipment at a 45-degree angle to the direction of laser beam propagation. The lens is equipped with an anti-reflective filter, and the sampling frequency is set to 500fps, outputting a color image sequence of 1280×720 pixels. It focuses on capturing the boundary morphology of the molten pool, the generation and movement trajectory of spatter particles. An acoustic sensor is installed at the bottom of the forming platform and fixed by an elastic bracket to reduce equipment vibration interference. The sampling frequency is set to 20kHz, and it outputs a one-dimensional time series with a length of 1024 sampling points to record the changes in sound waves during the solidification of the molten pool and the spreading of powder. A force sensor is integrated on the scraper connecting shaft, with a sampling frequency set to 50Hz, and it outputs a one-dimensional time series with a length of 512 sampling points to monitor the contact pressure fluctuations between the scraper and the surface of the formed part in real time. All sensors are connected to the data acquisition card via an industrial Ethernet interface. A timestamp synchronization protocol is used to control the timestamp error of each sensor's data within ±1ms, ensuring the spatiotemporal consistency of multi-source data.

[0057] Defect annotation employs a collaborative model combining manual and automatic annotation to ensure the comprehensiveness and accuracy of the results. In the manual annotation phase, a team of three technicians with over five years of experience in defect detection in metal additive manufacturing uses the LabelMe annotation tool to annotate the bounding boxes, defect types, and apparent dimensions of defect areas frame by frame, based on the acquired infrared temperature image sequences and molten pool morphology image sequences. If discrepancies exist in the annotation results among the three technicians (boundary box IoU < 0.7 or inconsistent defect types), a unified annotation conclusion is reached through collective discussion and analysis of image feature details. In the automatic annotation phase, for internal defects that are not observable manually, an industrial CT scanner is used to perform tomographic scanning of the formed part with a scanning accuracy of 0.05 mm. This acquires the three-dimensional structural data of the part's interior. A coordinate mapping algorithm converts the three-dimensional coordinates of the internal defects into image coordinates of the corresponding sensor data frames, supplementing the annotation with information on the location and size of internal pores and hidden cracks. Defect types are strictly enumerated into four categories: porosity, cracks, lack of fusion, and deformation. Porosity refers to internal voids formed during the solidification of the molten pool; cracks refer to linear cracks inside or on the surface of the formed part; lack of fusion refers to areas where adjacent layers or powder particles are not fully fused; and deformation refers to phenomena where the geometry of the formed part deviates from the design model. The annotation results are stored in XML format. Each annotation file contains information such as defect type, bounding box coordinates, dimensional parameters, annotator, and annotation time. A one-to-one association is established between timestamps and corresponding sensor data frames, ultimately forming the original dataset. Dataset The specific data items include: infrared temperature image sequences Molten pool morphology image sequence Acoustic signal sequence , contact force signal sequence Defect labeling information .

[0058] S2. Multi-source data preprocessing and fusion

[0059] To eliminate noise interference in the original data, standardize the data format, and extract effective features, the original dataset was modified. Each type of data undergoes specific preprocessing operations. All preprocessed data streams serve as the foundational input for subsequent multi-source fusion, ensuring the effectiveness of the fused features. The entire multi-source data processing workflow is as follows: Figure 2 As shown.

[0060] In the infrared temperature image preprocessing stage, median filtering is first used for noise removal. A 3×3 window size was chosen based on extensive experimental verification; this size effectively filters salt-and-pepper noise (generated by sensor electronic noise and ambient light interference) in the molten pool image while preserving the temperature gradient details at the molten pool edge to the greatest extent. Next, histogram equalization is performed, stretching the image's grayscale range to disperse temperature data originally concentrated in the low grayscale intervals to the full grayscale range, significantly improving the contrast between the high-temperature area in the center of the molten pool and the low-temperature area at the edge, facilitating subsequent feature extraction. Finally, bilinear interpolation is used to uniformly scale the image size to 256×256 pixels. This size balances feature preservation and model computation efficiency, resulting in the preprocessed infrared temperature image sequence. .

[0061] In the preprocessing stage of the molten pool morphology image, a Gaussian filtering algorithm is used to smooth noise, and the standard deviation is selected. A Gaussian kernel is used to effectively suppress Gaussian noise (caused by camera sensor noise) in the image by weighted averaging of the area surrounding each pixel. Since the molten pool region has significant characteristics in the saturation channel of the HSV color space, RGB-to-HSV color space conversion is used to extract the saturation channel image, highlighting the boundary difference between the molten pool region and the background powder, and reducing the impact of ambient light changes on the image. Similarly, a bilinear interpolation algorithm is used to scale the image to 256×256 pixels, resulting in a preprocessed sequence of molten pool morphology images. .

[0062] In the acoustic signal preprocessing stage, a wavelet thresholding denoising algorithm based on the db4 wavelet basis is used, with a decomposition level of 5. This wavelet basis has good time-frequency localization characteristics and can effectively separate environmental noise (such as equipment fan noise and workshop background noise) from defect-related signals (such as abrupt changes in sound waves caused by incomplete fusion) in the acoustic signal. The one-dimensional time-domain acoustic signal is then converted into a frequency-domain feature sequence using a Fast Fourier Transform (FFT), with the frequency range limited to 0-10kHz. This frequency band contains most of the defect-related acoustic feature information, resulting in a frequency-domain feature sequence with a dimension of 512. .

[0063] In the contact force signal preprocessing stage, a moving average filtering algorithm with a window size of 5 is used. By averaging five consecutive sampling points, high-frequency noise (generated by random friction between the scraper and powder) in the contact force signal is smoothed, while preserving the overall trend of the signal. A min-max normalization algorithm is then used to map the signal data to the [0,1] interval, eliminating the dimensional differences between the contact force signal and other sensor data, ensuring the fairness of the weights of each feature during the fusion process, and obtaining the preprocessed contact force signal sequence. The length remains 512.

[0064] Multi-source data fusion employs a feature-level fusion strategy. Its core objective is to integrate the spatial information of image features with the temporal / frequency domain information of one-dimensional signal features, forming a fused feature that combines multi-dimensional advantages. First, channel-level fusion is performed on image features. (Single channel) and (Single channel) Concatenate along the channel dimension, and calculate simultaneously. and The differential channel (reflecting the difference in changes between the two, highlighting the coordinated changes in molten pool temperature and morphology) and the product channel (reflecting the intensity correlation between the two, strengthening the correspondence between high-temperature regions and specific morphologies) ultimately form a 4-channel image feature. The dimensions are 256×256×4. Dimensional fusion is performed on one-dimensional signal features to obtain the frequency domain feature sequence. (Dimension 512) and contact force signal sequence (Dimension 512) are directly concatenated to form a one-dimensional fused feature vector with dimension 1024. .

[0065] To dynamically adjust the weight ratio of the two types of features, an attention mechanism is introduced, and the fusion formula is as follows:

[0066] ;

[0067] in, The attention weight is initially set to 0.6. This value is based on the statistics of historical defect detection data. Image features contribute slightly more to defect localization than one-dimensional signal features. It will be dynamically optimized through model training in the future. For the final fusion feature; 4-channel image features The flattened vector is obtained by unfolding the 256×256×4 feature map in row priority order, resulting in a one-dimensional vector with dimensions of 256×256×4=262144. One-dimensional fused feature vector The dimension mapping result is linearly mapped from a 1024-dimensional vector to a 262144-dimensional vector through a fully connected layer, ensuring consistency with... The dimensions are consistent. The fused features and their corresponding defect annotation information are consistent. One-to-one association to form a preprocessed dataset ,in The total number of data samples is denoted as , and each sample contains fused features and complete defect annotation information, providing standardized input for subsequent model training.

[0068] S3. Construction and Training of a Deep Learning Model Integrating Four Algorithms

[0069] The core of this step is to construct a deep learning model that integrates four algorithms: improved ResNet50 (CNN), multi-head self-attention Transformer, cross-attention mechanism, and Feature Pyramid Network (FPN). These four algorithms do not execute independently but form a deep collaborative system through bidirectional feature interaction and dynamic weight feedback. The core contribution lies in overcoming the limitations of single-algorithm feature extraction, while simultaneously capturing the local spatial details of defects, global correlation patterns, multi-scale feature differences, and cross-algorithm feature complementarity, providing high-precision and robust feature representation support for real-time defect detection. The construction and training process is as follows: Figure 3 As shown.

[0070] 3.1 Model Construction

[0071] 1. Algorithm 1: Improved ResNet50 (CNN)

[0072] The core function of the improved ResNet50 algorithm is to extract local spatial features of defects from the fused features, including detailed information such as defect boundary contours, local intensity distributions, and spatial relationships, providing fundamental feature support for subsequent defect localization. This algorithm uses preprocessed fused features... The input dimension is 256×256×4. Targeting the local feature scale of defects in metal additive manufacturing, six residual blocks are designed to form the main feature extraction body. Each residual block contains two 3×3 convolutional layers and one 1×1 shortcut connection. The kernel size of the 3×3 convolutional layers effectively captures the correlation information of local adjacent pixels, while the 1×1 shortcut connection addresses the gradient vanishing problem in deep networks, ensuring the effectiveness of feature extraction. Each convolutional layer is followed by a batch normalization layer and a ReLU activation function. The batch normalization layer standardizes the convolutional output, accelerating model training convergence, while the ReLU activation function introduces a non-linear transformation, enhancing the model's ability to fit complex defect features. Finally, a max-pooling layer with a stride of 2 downsamples the feature map, compressing the feature map size while retaining key features, outputting a local feature map. Its dimension is defined as ,in Indicates the feature map height. Indicates the width of the feature map. This represents the number of channels in the feature map. This dimension is designed to balance the preservation of details of local features with computational efficiency.

[0073] 2. Algorithm 2: Multi-head Self-Attention Transformer

[0074] The core function of the Transformer algorithm is to overcome the limitations of the local receptive field in CNNs, capturing the global correlations of defect features, such as crack propagation trends and pore distribution patterns, thus complementing the local features of CNNs. This algorithm will... Flattened into sequence features according to row priority. , dimension The hidden layer dimension is input to the Transformer encoder. The Transformer encoder contains 8 attention heads. This number, based on experimental verification, is sufficient to capture global correlations at different scales while avoiding a surge in computation caused by too many attention heads, ensuring real-time detection requirements are met; hidden layer dimension This provides sufficient dimensional space for feature transformation.

[0075] The association weights between sequence features are calculated using a self-attention mechanism. The self-attention formula is as follows:

[0076] ;

[0077] in, For querying the matrix, the dimension is , The key matrix has dimensions of . , It is a value matrix with dimension . All three are from It is generated through different linear transformations, and the weight matrix of the linear transformation is optimized through model training; For the sequence length, corresponding to The total number of pixels; As a feature dimension, ensure that the dimensions of the query, key, and value matrix match; This is a scaling factor used to mitigate... The problem of vanishing gradients in the Softmax function caused by excessively large matrix element values ​​is addressed. Each attention head independently computes its output features. The output features of the eight attention heads are concatenated along the channel dimension and then integrated into a unified feature map through a linear transformation. Its dimensions and Maintain consistency ( This ensures that the input dimensions match those of subsequent Cross-Attention.

[0078] 3. Algorithm 3: Cross-Attention Mechanism

[0079] The core function of the Cross-Attention algorithm is to achieve a two-way weighted fusion of local features extracted by CNNs and global features extracted by Transformers. By dynamically adjusting the weight ratio of the two types of features, it fully leverages the complementary advantages of local details and global correlations. Simultaneously, it provides optimized fused features to the FPN and receives feedback signals from the FPN to adjust the weights, forming a two-way interaction. This algorithm uses... and As input, the initial attention weights for the two types of features are first calculated. and The weight calculation formula is:

[0080] ;

[0081] ;

[0082] in, for The initial weights, for The initial weights; for The global average eigenvalue is calculated as follows: , express The eigenvalues ​​at position (h, w) and channel c; for The global average eigenvalue, calculated in the same way as Consistent, express The eigenvalues ​​at position (h, w) and channel c; the Softmax function is used to convert the global average eigenvalues ​​into a probability distribution, reflecting the initial importance ratio of the two types of features.

[0083] Preliminary fusion features are obtained based on the initial weights. The calculation formula is:

[0084] ;

[0085] Dimensional preservation This feature integrates local details and global correlations, but the weights do not yet take into account the influence of multi-scale features and need to be optimized through FPN feedback.

[0086] 4. Algorithm 4: Feature Pyramid Network (FPN)

[0087] The core function of the FPN algorithm is to extract multi-scale features of defects, adapting to the detection needs of defects of different sizes (such as micropores and large-area deformations). Simultaneously, through information entropy analysis of multi-scale features, it feeds back adjustment signals to the Cross-Attention mechanism, optimizing the fusion weights of local and global features, forming a two-way interactive closed loop between algorithms. This algorithm will... Input three concatenated convolutional blocks to generate feature maps at different scales:

[0088] First-scale feature map Processed using a single 3×3 convolutional layer (stride 1). The number of convolutional kernels is set to 512, and the output dimension is... ,in , , This scale characteristic corresponds to medium-sized defects (such as pores of 0.1-0.5 mm).

[0089] Second-scale feature map :right Perform a 3×3 convolution with a stride of 2, set the number of kernels to 1024, and output dimension... ,in , , This dimensional characteristic corresponds to larger size defects (such as cracks of 0.5-2 mm).

[0090] Third-scale feature map :right Perform a 3×3 convolution with a stride of 2, set the number of kernels to 2048, and output dimension... ,in , , This dimensional feature corresponds to ultra-large defects (such as unfused or deformed defects larger than 2 mm).

[0091] To integrate the advantages of multi-scale features, an upsampling and stitching strategy is adopted: for Upsampling using bilinear interpolation with an interpolation factor of 2, scaled to Size, and Concatenate according to channel dimension; for Upsampling using bilinear interpolation with an interpolation factor of 4, scaled to The dimensions are then combined with the above-mentioned stitching results again along the channel dimension to obtain multi-scale fused features. , dimension This feature contains characteristic information about defects of different sizes, which significantly improves the model's ability to adapt to defects of multiple sizes.

[0092] To achieve bidirectional interaction with Cross-Attention, FPN calculates the feedback adjustment coefficients using the information entropy of multi-scale features. Information entropy reflects the richness of information in a feature; the higher the information entropy, the richer the defect information contained in the feature. The formula for calculating information entropy is:

[0093] ;

[0094] ;

[0095] in, Feature maps corresponding to three scales respectively; The information entropy of the feature map at the k-th scale; Eigenvalues The probability distribution is obtained through normalization. The base is logarithmic, ensuring that the unit of information entropy is bits.

[0096] Feedback adjustment coefficient The calculation formula is:

[0097] ;

[0098] The value range is [0,1], reflecting the balance of information richness of the feature maps at the three scales: The closer it is to 1, the more balanced the information richness of the multi-scale features and the better the multi-scale fusion effect. The closer it is to 0, the higher the information content of a certain scale feature is, while other scale features contribute insufficiently.

[0099] Will Feedback is fed into Cross-Attention to dynamically adjust the initial weights. and The formula is adjusted as follows:

[0100] ;

[0101] ;

[0102] ;

[0103] ;

[0104] in, , The adjusted intermediate weights, The larger, The greater the increase, the stronger the weight of local features; The smaller, The greater the increase, the stronger the weight of the global features, thus achieving dynamic weight optimization based on multi-scale feature feedback; , The final weights after normalization are set to ensure that the sum of the weights is 1, which meets the requirements of the probability distribution.

[0105] The optimized fusion features are obtained based on the final weights. The calculation formula is:

[0106] ;

[0107] Dimensional preservation This feature contains both precise local details and integrates global correlation information, and its weights are optimized through multi-scale feature feedback, giving it a stronger defect representation capability.

[0108] Will The final fusion layer of the input FPN is compressed to 256 channels using a 1×1 convolutional layer to obtain the core feature map output by the model. , dimension This feature will be directly used for subsequent defect localization and classification tasks, and is the final feature output of the four types of algorithms working together.

[0109] 3.2 Model Training Process

[0110] The core objective of model training is to optimize all parameters of the four types of algorithms to make the core feature maps... It can accurately represent defect information, ensuring the accuracy of defect location, classification and severity assessment. The training process strictly follows the process of data partitioning, parameter initialization, hyperparameter setting, loss calculation and iterative optimization to ensure training effect.

[0111] 1. Dataset partitioning: Divide the preprocessed dataset into parts. The dataset was randomly divided into training sets in a 7:2:1 ratio. Validation set Test set Training set Used for iterative updates of model parameters, containing 70% of the samples; validation set. Used to monitor overfitting during model training and adjust hyperparameters, including 20% ​​of the samples; test set. Used to evaluate the final performance of the model after training, it contains 10% of the samples, and the samples from the three datasets do not overlap to ensure the objectivity of the evaluation results.

[0112] 2. Parameter Initialization: Different initialization strategies are adopted for the parameter characteristics of different algorithms. The parameters of the CNN layers of the improved ResNet50 are initialized using the He normal distribution. This initialization method is based on the characteristics of the ReLU activation function, which can keep the variance of the output of each layer consistent and accelerate training convergence. The parameters of the Transformer layer and the Cross-Attention layer are initialized using the Xavier uniform distribution. This method ensures that the variance of the input and output is consistent and avoids gradient vanishing or exploding. The convolution kernel parameters of the FPN layer are initialized with a constant initial value of 0.01 to ensure gradient stability in the initial training phase.

[0113] 3. Hyperparameter settings: Batch size is set to 32, a value determined based on GPU memory capacity (16GB), balancing training efficiency and memory usage; training epochs are set to 100 to ensure the model has sufficient iterations to fit complex defect features; initial learning rate is set to... A cosine annealing learning rate scheduling strategy is adopted, which decays the learning rate to 0.5 times the current value every 20 epochs. This strategy can quickly explore the parameter space in the early stage of training and accurately adjust the parameters in the later stage, thereby improving the model's generalization ability. The optimizer used is AdamW, and the weight decay coefficient is set to L2 regularization is used to suppress overfitting and improve the robustness of the model.

[0114] 4. Loss Function Design: A multi-task loss fusion strategy is adopted, comprehensively considering the optimization objectives of defect localization, classification, and feature fusion. The loss function formula is as follows:

[0115] ;

[0116] in, , , To determine the weighting of the loss, it is set based on the importance ratio of the three types of tasks. Positioning accuracy is the foundation of defect detection, so it is given the highest weight. The GIoU loss is used to optimize the positional deviation between the predicted defect bounding box and the actual labeled bounding box. The calculation formula is as follows: , For the defect boxes predicted by the model, For the truly labeled defect boxes, GIoU simultaneously considers the overlapping areas and the area difference between the predicted box and the true box, resulting in better localization accuracy than traditional IoU loss. Cross-entropy loss is used to optimize the probability prediction accuracy of defect classification, and its calculation formula is as follows: , The unique thermal encoding label for the defect type (e.g., the pore corresponds to [1,0,0,0]). The probability of various defects predicted by the model; To fuse feature loss, used for optimization and Feature similarity is used to ensure the effectiveness of bidirectional fusion, and the calculation formula is as follows: ,in , , for Dimensions for Feature values ​​after downsampling to a size of 64×64.

[0117] 5. Training and Optimization: [The text abruptly ends here, likely due to an incomplete sentence or a format Fusion features in The model is input in batches and then processed sequentially using an improved ResNet50 to generate... Transformer generation Cross-Attention generation FPN generation to And calculate Adjusting the weights to obtain Final generation ;based on Predicted defect boxes and classification probability (j=1-4), through The batch loss is calculated, the gradients of each parameter are computed using the backpropagation algorithm, and all parameters of the four algorithms are updated using the AdamW optimizer. Every 5 epochs... To validate the model performance, calculate the localization accuracy (the proportion of predicted boxes with IoU ≥ 0.5), classification accuracy (the proportion of correctly classified defective samples), and fusion feature similarity (…). and (cosine similarity); when the validation set performance shows no improvement for 10 consecutive epochs, an early stopping strategy is triggered to stop training and save the current optimal model parameters. To avoid overfitting, use the test set. right Perform a final performance evaluation. If the classification accuracy is below 95% or the localization accuracy is below 90%, adjust the algorithm hyperparameters (e.g., increase the number of Transformer attention heads to 12 and the number of FPN scales to 4), and re-execute the training process until the model performance meets the requirements.

[0118] 3.3 Model Application Process

[0119] The core of the model application phase is to optimize the trained model. Deployed to edge computing nodes, enabling rapid real-time data processing and defect detection. After the metal additive manufacturing equipment is started, S4 outputs real-time fused features. Will continue to input The model performs inference according to the following process: First, it extracts real-time local features by improving ResNet50. It captures local details of defects in the current frame; then it extracts real-time global features using a Transformer. Analyze the global correlation information of defects; based on and Calculate real-time initial weights and Generate preliminary fusion features ;Will Input FPN to generate real-time multi-scale features at three scales. , , Calculate the real-time feedback coefficient The final weight is obtained by dynamic adjustment. and Output optimized real-time fusion features Finally, the real-time core feature map is generated through the FPN final fusion layer. This feature map is directly passed to the defect localization process in S5 and the defect classification process in S6, providing core support for real-time detection. The entire inference process strictly follows the algorithm logic of the training phase to ensure the consistency between the real-time detection results and the training effect, while optimizing the inference speed, with the single-frame processing time controlled within 50ms, meeting the real-time detection requirements of metal additive manufacturing.

[0120] S4. Real-time data acquisition and online preprocessing

[0121] After the metal additive manufacturing equipment is started, the four types of sensors deployed in S1 begin real-time data acquisition according to preset parameters, ensuring the spatiotemporal consistency and parameter stability of the acquired data. Infrared thermal imager, high-speed camera, acoustic sensor, and force sensor simultaneously acquire infrared temperature images. Molten pool morphology images Acoustic signals , contact force signal Data is transmitted to the edge computing node in the form of a data stream through an industrial Ethernet interface. The transmission protocol adopts TCP / IP to ensure the reliability of data transmission. The hardware configuration of the edge computing node is CPU (Intel Core i7-12700H) and GPU (NVIDIA RTX 3070Ti), which can meet the computing requirements of real-time preprocessing and model inference. The data transmission latency is controlled within ≤50ms to avoid detection delays caused by latency.

[0122] Edge computing nodes pre-store preprocessing parameter configuration files, which are completely identical to the parameters used in S2 offline preprocessing, ensuring data format consistency between real-time preprocessing and offline training. The online preprocessing process strictly follows the logic of S2: for A 3×3 windowed mid-range filter was performed to remove real-time salt-and-pepper noise. Histogram equalization was used to improve temperature contrast, and then bilinear interpolation was used to scale the image to 256×256 pixels. ;right implement After Gaussian filtering, RGB to HSV color space conversion, saturation channel extraction, and scaling to 256×256 pixels, the result is... ;right Wavelet thresholding with a db4 wavelet basis and 5-level decomposition is used for denoising, followed by Fast Fourier Transform to convert it into a frequency domain feature sequence of 0-10kHz, resulting in... (Dimension 512); to Using a 5-window moving average filter, min-max is normalized to the [0,1] interval to obtain (Length 512).

[0123] Online multi-source data fusion strictly follows the S2 fusion formula and generates fusion features in real time. First of all and stitched together as 4-channel image features flattened (262144 dimensions); will and Concatenate into a one-dimensional feature vector (1024-dimensional), mapped to 262144-dimensional through a fully connected layer, resulting in Attention weights optimized during the training phase (stored in) (in the parameter file), calculate . Once generated, it is immediately transmitted to the S5 real-time defect detection process as... Input features.

[0124] S5. Real-time Defect Detection and Location

[0125] Online preprocessing generated Continuously input the optimal model The model uses S3's four-category fusion algorithm for fast inference and outputs real-time core feature maps. ,based on Perform defect location detection. An anchor frame mechanism is used, pre-defined... Anchor frames of three sizes (16×16, 32×32, 64×64 pixels) and three aspect ratios (1:1, 1:2, 2:1) are set to cover defects of different sizes and shapes; convolutional layers are used to... Perform feature mapping to predict the defect confidence score for each anchor box. And coordinate offset, which is used to correct the anchor frame position to obtain the final defect prediction box coordinates. ,in , The coordinates of the top left corner of the prediction box. , The coordinates are the bottom right corner of the prediction box.

[0126] Set confidence threshold This threshold is determined statistically from the training set, specifically the minimum confidence level of all correctly located defect prediction boxes in the training set, ensuring that the threshold effectively filters out false positive detection results. If a defect is found in the current frame, the image coordinates of the prediction box need to be converted to the physical coordinates of the device's forming area. This coordinate transformation is achieved through a calibration matrix between the device coordinate system and the image coordinate system. Pre-calibration is performed using a standard calibration plate, and the physical coordinates of the standard points on the calibration plate in the equipment coordinate system are determined. pixel coordinates in the image coordinate system Given that the solution obtained by the least squares method is... The conversion formula is:

[0127] ;

[0128] in , The image coordinates of the center of the prediction box. , , The physical coordinates of the defect center. The coordinates are determined based on the current forming layer thickness information, with a calibration accuracy of ±0.1mm, ensuring precise location of defects. If... If the current frame is found to be defect-free, the edge computing node continues to receive the next batch of real-time data, returns to S4 to perform online preprocessing, and forms a continuous detection loop.

[0129] S6. Defect Classification and Severity Assessment

[0130] In cases where S5 determination has defects, based on Output Defect classification and severity assessment are performed simultaneously. The defect classification process will... Flattened into a one-dimensional vector in row-major order, with dimensions of 64×64×256=1048576, this vector is input to two fully connected layers: the hidden layer of the first fully connected layer has a hidden layer dimension of 1024 and uses the ReLU activation function for feature dimension compression and nonlinear transformation; the output dimension of the second fully connected layer is set to 4, corresponding to the four defect types, and the output is converted into a probability distribution using the Softmax activation function. ,in This represents the predicted probability of pore defects. This represents the predicted probability of a crack defect. This represents the predicted probability of a non-fusion defect. This represents the predicted probability of a deformation defect. The defect type corresponding to the highest probability is taken as the final classification result. For example, if... If the value is the maximum, then the current defect is determined to be a crack defect.

[0131] Severity assessment process based on The system outputs a severity score based on the characteristic intensity and defect size information. (Value range 0-1), the higher the score, the greater the impact of the defect on the performance of the formed part. The calculation is based on the mapping relationship between defect features and severity learned during the model training phase. Specifically, the model establishes a non-linear mapping between feature strength and severity by learning information such as defect size (area, length, depth), defect location (whether it is located in a critical stress area), and defect type labeled in the training set. Three severity thresholds are preset: , This threshold is determined through mechanical property tests on molded parts: tensile and bending tests are conducted on molded parts containing defects of varying severity, and the relationship between defect severity and mechanical properties (tensile strength, yield strength) is statistically analyzed to ultimately determine the threshold classification criteria. If the defect is small, it is considered a minor defect. Such defects have little impact on the mechanical properties of the formed part and can be suppressed through subsequent process adjustments. If the defect is not severe, it is classified as a moderate defect, requiring equipment shutdown for parameter adjustments to prevent the defect from worsening; if... If the defect is found to be severe, it may lead to the scrapping of the molded part. Operation must be stopped immediately and the fault investigated.

[0132] Defect type, physical location The severity level and testing timestamp are integrated into the testing report. The detection timestamp is consistent with the timestamp of the sensor data, which facilitates subsequent tracing of the specific process stage at which the defect occurred. The data is transmitted in real time via industrial Ethernet to the control system and host computer monitoring platform of the metal additive manufacturing equipment. The control system is used to execute subsequent process adjustments, while the host computer monitoring platform is used to display defect information in real time for operators to view.

[0133] S7. Feedback Adjustment and Dataset Update

[0134] The control system of the metal additive manufacturing equipment receives the test report. Then, based on the severity level of the defect, a targeted process parameter adjustment strategy is implemented to achieve real-time suppression and closed-loop control of the defect.

[0135] For minor defects, the control system does not interrupt equipment operation and adjusts the corresponding process parameters based on the defect type: if the defect is a pore, the laser power is increased by 5% to promote bubble overflow by increasing the molten pool temperature; if the defect is a crack, the scanning speed is decreased by 3% to prolong the molten pool solidification time and reduce internal stress; if the defect is incomplete fusion, the laser power is increased by 3% while keeping the scanning speed constant to enhance powder fusion; if the defect is deformation, the layer thickness is decreased by 2% to reduce the cumulative stress of the formed part. After adjustment, the defect detection results of subsequent frames are continuously monitored to observe whether the defects are suppressed.

[0136] For moderate defects, the control system pauses equipment operation for 10 seconds and performs multi-parameter coordinated adjustments: if the defect is a pore, the laser power is increased by 10% and the scanning interval is decreased by 0.05 mm to improve the spread and temperature of the molten pool; if the defect is a crack, the laser power is increased by 8%, the scanning speed is decreased by 5%, and the layer thickness is decreased by 0.02 mm to comprehensively reduce internal stress and solidification rate; if the defect is incomplete fusion, the laser power is increased by 10%, the scanning interval is decreased by 0.05 mm, and the scanning speed is decreased by 3% to enhance interlayer and interparticle fusion; if the defect is deformation, the scanning speed is decreased by 5%, the layer thickness is decreased by 0.02 mm, and the laser power is homogenized (power is reduced by 3% in the edge area) to balance the temperature field distribution of the formed part. After the adjustment is completed, the equipment operation is resumed, and the defect changes are continuously monitored.

[0137] For severe defects, the control system immediately shuts down the equipment, and the host computer monitoring platform issues an audible and visual alarm signal. The alarm sound has a frequency of 1kHz, and the alarm light flashes red, prompting the operator to handle the situation promptly. The operator needs to check the equipment parameters (such as laser power stability and scanning path accuracy), raw material quality (such as particle size distribution and purity of metal powder), and equipment status (such as scraper wear and forming platform levelness). After troubleshooting, the equipment should be restarted to prevent the recurrence of severe defects.

[0138] While performing feedback adjustments, the edge computing nodes automatically feed back real-time detection-related data to the original dataset of S1, enabling dynamic updates to the dataset. Updated data items include: real-time fused features. Real-time core feature map Test report Adjusted process parameters (such as laser power, scanning speed, layer thickness, and scanning spacing). Store this data in S1 format, along with the original dataset. Merge to form an updated dataset ,in This is the adjusted set of process parameters. When... The number of samples is greater than that of the original dataset. When the value increases by 10%, the incremental training process of the model is triggered, and the S3 pair is returned. Perform incremental training: maintain batch size at 32, set training epochs to 20, and initial learning rate to [value missing]. (A learning rate lower than the initial training rate to avoid overwriting existing effective parameters), the optimizer still uses AdamW, with a weight decay coefficient of [value missing]. Through incremental training, the model can learn new defect features and feature changes after process adjustments, continuously improving detection accuracy and adaptability, forming a closed-loop system of data acquisition, model training, real-time detection, feedback adjustment, data update, and model optimization. The real-time detection and closed-loop optimization processes are as follows: Figure 4 As shown.

[0139] In one specific embodiment:

[0140] I. Implementation Scenario Description

[0141] This embodiment describes a scenario where a titanium alloy Ti6Al4V aero-engine blade is manufactured using selective laser melting (SLM) equipment. The blade dimensions are 100mm × 30mm × 5mm, with a forming accuracy requirement of ±0.05mm. The defect tolerances are: pore diameter ≤ 0.1mm, crack length ≤ 0.2mm, and unfused area ≤ 0.5mm. For blades with deformation ≤0.1mm, four types of defects need to be detected and suppressed in real time to ensure the mechanical properties of the blades (tensile strength ≥900MPa, yield strength ≥820MPa).

[0142] II. Implementation Preparation

[0143] 1. Equipment and sensor configuration

[0144] SLM equipment: Farsoon FS271M, laser power range 50-500W, scanning speed 0-2000mm / s, layer thickness range 0.02-0.1mm;

[0145] Infrared thermal imager: FLIR A655sc, temperature measurement range -40-1500℃, 640×480 pixels, 100Hz frame rate;

[0146] High-speed camera: Phantom V2512, 1280×720 pixels, 500fps frame rate, equipped with an 850nm anti-reflective filter;

[0147] Acoustic sensor: PCB Piezotronics 356A16, sensitivity 10mV / m / Sampling frequency 20kHz;

[0148] Force sensor: HBM U9C, range 0-50N, accuracy ±0.1%FS, sampling frequency 50Hz;

[0149] Edge computing node: Intel Core i7-12700H CPU + NVIDIA RTX 3070Ti GPU (16GB VRAM), 32GB RAM.

[0150] 2. Materials and Process Parameters

[0151] Metal powder: Ti6Al4V spherical powder, particle size 15-53μm, purity ≥99.9%;

[0152] Basic process parameters: laser power 280W, scanning speed 1200mm / s, layer thickness 0.04mm, scanning interval 0.12mm, forming chamber argon atmosphere (oxygen content ≤0.1%).

[0153] 3. Dataset Preparation

[0154] Multi-source data was collected from the forming process of 100 blades, each containing 12,500 frames of sensor data (forming 500 layers, 25 frames per layer), totaling 1.25 × Frame data;

[0155] Manual annotation + industrial CT annotation: 3200 pores, 1800 cracks, 1500 lack of fusion, and 900 deformations were annotated to form the original dataset. The training set was divided into two parts in a 7:2:1 ratio (8.75× Frames), validation set (2.5× Frames), test set (1.25× frame).

[0156] III. Detailed Implementation Steps

[0157] S1. Multi-source sensor data acquisition and defect labeling

[0158] Sensor deployment: An infrared thermal imager is installed directly above the equipment's observation window (300mm from the molten pool), a high-speed camera is deployed on the side of the equipment (at a 45° angle to the laser beam, 250mm from the molten pool), an acoustic sensor is fixed to the bottom of the forming platform, and a force sensor is integrated into the scraper connecting shaft;

[0159] Synchronous acquisition: Data is transmitted via industrial Ethernet TCP / IP protocol, with timestamp synchronization error controlled within ±0.8ms. During the acquisition process, process parameters such as layer number, laser power, and scanning speed of each frame of data are recorded.

[0160] Defect labeling: Three senior technicians used LabelMe to label surface defects, and the industrial CT (ZEISS Metrotom1500, scanning accuracy 0.05mm) was used to label internal defects. The labeling files were stored in XML format and were linked to the sensor data through timestamps and layer numbers.

[0161] S2. Multi-source data preprocessing and fusion

[0162] Infrared temperature image preprocessing: 3×3 median filtering → histogram equalization → bilinear interpolation scaling to 256×256, resulting in... ;

[0163] Preprocessing of molten pool morphology images: Gaussian filtering → RGB to HSV conversion, extract saturation channel → scale to 256×256, resulting in... ;

[0164] Acoustic signal preprocessing: 5-layer decomposition and denoising using a db4 wavelet basis → FFT conversion to 0-10kHz frequency domain features (dimension 512), resulting in... ;

[0165] Contact force signal preprocessing: 5-window moving average filtering → min-max normalization to [0,1], resulting in ;

[0166] Feature fusion: and stitched together as 4-channel image features (256×256×4), flattened to (262144 dimensions); and spliced ​​as (1024-dimensional), mapped to 262144-dimensional through a fully connected layer to obtain Attention weights Initial value 0.6, generate fusion features according to the formula:

[0167] ;

[0168] Fusion features and annotation information Association formation .

[0169] S3. Construction and Training of a Deep Learning Model Integrating Four Algorithms

[0170] 3.1 Model Construction

[0171] Improved ResNet50: 6 residual blocks, input (256×256×4), Output (64×64×256);

[0172] Transformer encoder: 8 attention heads, 512 hidden layer dimensions, input... (4096×256), Output (64×64×256), Self-attention calculation:

[0173] ;

[0174] Cross-Attention: Calculate the initial weights:

[0175] ;

[0176] ;

[0177] generate (64×64×256):

[0178] ;

[0179] FPN: Generation (32×32×512) (16×16×1024) (8×8×2048), obtained by upsampling and stitching (32×32×3584); Calculate information entropy:

[0180] ;

[0181] ;

[0182] Feedback adjustment coefficient:

[0183] ;

[0184] Adjusting weights:

[0185] ;

[0186] ;

[0187] ;

[0188] ;

[0189] generate (64×64×256):

[0190] ;

[0191] Final output (64×64×256).

[0192] 3.2 Model Training

[0193] Hyperparameters: Batch size 32, number of training epochs 100, initial learning rate Cosine annealing scheduling (decays by 50% every 20 epochs), AdamW optimizer (weight decay). );

[0194] Loss function:

[0195] ;

[0196] in , , ;

[0197] Training results: Validation set localization accuracy 93.2%, classification accuracy 96.5%, fusion feature similarity 0.92. Early stopping was triggered after 58 epochs of training, and the results were saved. The test set achieved a localization accuracy of 92.8%, a classification accuracy of 95.7%, and a severity prediction MAE of 0.03.

[0198] 3.3 Model Application Deployment

[0199] Will Deployed to edge computing nodes, the inference engine is accelerated using TensorRT, with a single frame processing time of 38ms, meeting the requirements for real-time detection.

[0200] S4. Real-time data acquisition and online preprocessing

[0201] After the SLM device is started, the four types of sensors synchronously collect data according to preset parameters. , , , Data transmission delay is 42ms;

[0202] Online preprocessing of edge computing nodes: generated according to the S2 process , , , Real-time fusion The data is then transmitted to the defect detection process.

[0203] S5. Real-time Defect Detection and Location

[0204] enter Output Predicting defect box coordinates and confidence levels using the anchor box mechanism. ;

[0205] set up A certain frame detected The system was found to have a defect; the predicted bounding box coordinates (120, 85, 145, 110) were converted to physical coordinates using a calibration matrix.

[0206] ;

[0207] Corresponding to layer 310, the positioning accuracy is ±0.08mm.

[0208] S6. Defect Classification and Severity Assessment

[0209] Output probability distribution It was determined to be a crack defect;

[0210] Severity score , in and Between these ranges, it is determined to be a moderate defect.

[0211] S7. Feedback Adjustment and Dataset Update

[0212] The control system paused the equipment operation for 10 seconds and implemented a moderate crack adjustment strategy: laser power increased by 8% (302.4W), scanning speed decreased by 5% (1140mm / s), and layer thickness decreased by 0.02mm (0.038mm).

[0213] After adjustments, the system resumed operation, and no further crack defects were detected in the subsequent 5 frames.

[0214] Edge computing nodes will , , The adjusted process parameters are updated to the dataset. ,form ;

[0215] when When the number of samples increases by 10%, return to S3 for incremental training (20 epochs, learning rate). After incremental training, the classification accuracy on the test set improved to 96.3%.

[0216] IV. Implementation Results

[0217] Defect detection performance: Pore detection rate 94.1%, crack detection rate 95.8%, incomplete fusion detection rate 93.5%, deformation detection rate 92.3%, average detection delay 45ms;

[0218] Forming quality: 100 blades were inspected by X-ray flaw detection, and the defect rate was reduced from 12% by traditional methods to 2.1%. The mechanical properties all met the requirements (tensile strength 915-930MPa, yield strength 830-845MPa).

[0219] Production efficiency: The scrap rate due to defects was reduced from 8% to 1.2%, and the production cycle of a single blade was shortened by about 15% (reducing rework time).

[0220] The embodiments described above are for illustrative purposes only and are not intended to limit the invention. Therefore, any changes in numerical values ​​or substitutions of equivalent elements should still fall within the scope of this invention.

[0221] The above detailed description will enable those skilled in the art to understand that the present invention can indeed achieve the aforementioned objectives and has complied with the provisions of the Patent Law.

[0222] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention. The above descriptions are merely preferred embodiments of the invention and are not intended to limit the invention. It should be noted that any modifications, equivalent substitutions, and improvements made within the spirit and principles of the invention should be included within the scope of protection of the invention.

[0223] It should be noted that the above description of the process is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art can make various modifications and changes to the process under the guidance of this specification. However, these modifications and changes remain within the scope of this specification.

[0224] The basic concepts have been described above. Obviously, for those skilled in the art who have read this application, the above disclosure is merely illustrative and does not constitute a limitation of this application. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this application. Such modifications, improvements, and corrections are suggested in this application, and therefore, such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of this application.

[0225] Furthermore, this application uses specific terms to describe its embodiments. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different positions in this specification do not necessarily refer to the same embodiment. In addition, certain features, structures, or characteristics in one or more embodiments of this application can be appropriately combined.

[0226] Furthermore, those skilled in the art will understand that aspects of this application can be described and illustrated through several patentable types or situations, including any new and useful combination of processes, machines, products, or substances, or any new and useful improvements thereof. Therefore, aspects of this application can be implemented entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or a combination of hardware and software. All of the above hardware or software can be referred to as a “unit,” “module,” or “system.” Furthermore, aspects of this application can take the form of a computer program product embodied in one or more computer-readable media, wherein computer-readable program code is contained therein.

[0227] The computer program code required for the operation of each part of this application can be written in any one or more programming languages, including object-oriented programming languages ​​such as Java, Scala, Smalltalk, Eiffel, JADE, Emerald, C++, C#, VB.NET, and Python; general programming languages ​​such as C; Visual Basic, Fortran2103, Perl, COBOL2102, PHP, and ABAP; dynamic programming languages ​​such as Python, Ruby, and Groovy; or other programming languages. This program code can run entirely on the user's computer, or as a standalone software package on the user's computer, or partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any network, such as a local area network (LAN) or wide area network (WAN), or connected to an external computer (e.g., via the Internet), or in a cloud computing environment, or used as a service such as Software as a Service (SaaS).

[0228] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this application are not intended to limit the order of the processes and methods of this application. Although some currently considered useful embodiments of the invention have been discussed in the foregoing disclosure by way of various examples, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the substance and scope of the embodiments of this application. For example, although the implementation of the various components described above can be embodied in a hardware device, it can also be implemented as a purely software solution, such as an installation on an existing server or mobile device.

[0229] Similarly, it should be noted that, in order to simplify the description of the present application and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of the embodiments of the present application sometimes combines multiple features into a single embodiment, drawing, or description thereof. However, this approach of the present application should not be construed as reflecting an intention that the claimed subject matter requires more features than expressly recited in each claim. Rather, the subject of the invention should possess fewer features than in any single embodiment described above.

Claims

1. A deep learning method for real-time defect detection in metal additive manufacturing, characterized in that, include: S1. Simultaneously acquire infrared temperature image sequences, molten pool morphology image sequences, acoustic signal sequences, and contact force signal sequences during the metal additive manufacturing process, and perform defect annotation to form an original dataset containing the acquired data and defect annotation information; S2. Targeted preprocessing is performed on the collected data in the original dataset. The preprocessed infrared temperature image sequence and the melt pool morphology image sequence are fused at the channel level, and the preprocessed acoustic signal sequence and the contact force signal sequence are fused at the dimension level. An attention mechanism is introduced to dynamically adjust the weight ratio of the two types of fused features to generate a unified fused feature. This fused feature is associated with the defect annotation information in the original dataset to form a preprocessed dataset. The channel-level fusion forms a 4-channel image feature, which specifically includes: the infrared temperature image preprocessing result channel, the melt pool morphology image preprocessing result channel, the difference channel between the two, and the product channel between the two. S3. Construct a deep learning model that integrates an improved ResNet50, a multi-head self-attention Transformer, a cross-attention mechanism, and a feature pyramid network. The improved ResNet50 is used to extract local spatial features from the fused features. The Transformer is used to extract global correlation features from the fused features. The cross-attention mechanism is used to dynamically fuse local spatial features and global correlation features and output preliminary fused features. The feature pyramid network is used to extract multi-scale features from the preliminary fused features and generate feedback adjustment coefficients through information entropy analysis. The cross-attention mechanism receives the feedback adjustment coefficients, optimizes the fusion weights, and outputs the final fused features. Train the deep learning model based on the preprocessed dataset to obtain the optimal model. Feedback adjustment coefficient The calculation formula is: ; The information entropy of the feature maps at three different scales generated by the feature pyramid network are respectively; The deep learning model is trained using a multi-task loss function: ;in, , , To lose weight, The GIoU loss is used to optimize the positional deviation between the predicted defect bounding box and the actual labeled bounding box. Cross-entropy loss is used to optimize the accuracy of probability prediction for defect classification. The fusion feature loss is used to optimize the similarity between the final fused features output by the cross-attention mechanism and the multi-scale fused features output by the feature pyramid network. S4. During the operation of the metal additive manufacturing equipment, real-time infrared temperature image sequence, real-time molten pool morphology image sequence, real-time acoustic signal sequence, and real-time contact force signal sequence are collected simultaneously. The real-time collected data are processed online according to the preprocessing process in step S2, and then real-time fusion features are generated through the fusion method in step S2. S5. Input the real-time fused features into the optimal model, and output the real-time core feature map through collaborative reasoning of the four types of algorithms in the optimal model. Based on the real-time core feature map, use the anchor box mechanism to predict the defect location and confidence level. When the confidence level reaches the preset threshold, convert the image coordinates of the defect into the physical coordinates of the equipment forming area to determine that a defect exists.

2. The deep learning method for real-time defect detection in metal additive manufacturing according to claim 1, characterized in that, Also includes: S6. Defect Classification and Severity Assessment: Based on the real-time core feature map output by the optimal model, the probability distribution of various defects is obtained through fully connected layer mapping to determine the defect type. At the same time, the severity score is output by combining the defect feature intensity and size information. According to the preset threshold, the defects are divided into three levels: mild, moderate and severe. The defect type, physical coordinates, severity level and detection timestamp are integrated to form a detection report.

3. The deep learning method for real-time defect detection in metal additive manufacturing according to claim 2, characterized in that, Also includes: S7. Feedback Adjustment and Dataset Update: The equipment control system receives the inspection report, executes the corresponding process parameter adjustment strategy according to the defect level, and feeds back the real-time fused features, real-time core feature map, inspection report and adjusted process parameters to the original dataset in step S1 to form the updated dataset. When the number of samples in the updated dataset increases by a preset ratio compared with the original dataset, it returns to step S3 to incrementally train the optimal model to form a closed-loop optimization.

4. The deep learning method for real-time defect detection in metal additive manufacturing according to claim 1, characterized in that, In step S1, infrared temperature image sequence, molten pool morphology image sequence, acoustic signal sequence, and contact force signal sequence are simultaneously acquired by infrared thermal imager, high-speed camera, acoustic sensor, and force sensor during the metal additive manufacturing process. The four types of sensors transmit data through an industrial Ethernet interface using a timestamp synchronization protocol, with the timestamp error controlled within ±1ms. The types of defects marked include pores, cracks, lack of fusion, and deformation. The marking results are linked one-to-one with the sensor data through timestamps.

5. The deep learning method for real-time defect detection in metal additive manufacturing according to claim 1, characterized in that, In step S2, the infrared temperature image preprocessing includes 3×3 window midpoint filtering, histogram equalization, and bilinear interpolation scaling; the melt pool morphology image preprocessing includes Gaussian filtering with a standard deviation of 1.5, RGB to HSV color space conversion, and bilinear interpolation scaling; the acoustic signal preprocessing includes db4 wavelet basis 5-layer decomposition wavelet thresholding denoising and fast Fourier transform frequency domain conversion; and the contact force signal preprocessing includes 5-window moving average filtering and min-max normalization.

6. The deep learning method for real-time defect detection in metal additive manufacturing according to claim 1, characterized in that, In step S2, channel-level fusion forms 4-channel image features, including the infrared temperature image preprocessing result channel, the molten pool morphology image preprocessing result channel, the difference channel and the product channel of the two. One-dimensional signal data is fused into a one-dimensional feature vector, which is mapped to the same dimension as the flattened 4-channel image features through a fully connected layer, and then participates in attention-weighted fusion.

7. The deep learning method for real-time defect detection in metal additive manufacturing according to claim 1, characterized in that, In step S3, the improved ResNet50 contains 6 residual blocks. Each residual block consists of two 3×3 convolutional layers connected to one 1×1 shortcut. Each convolutional layer is followed by a batch normalization layer and a ReLU activation function. Finally, a local feature map is output through a max pooling layer with a stride of 2.

8. The deep learning method for real-time defect detection in metal additive manufacturing according to claim 1, characterized in that, In step S3, the cross-attention mechanism first calculates the initial weights based on the global average feature value of the local feature map and the global associated feature map, and then dynamically adjusts the weights by receiving the feedback adjustment coefficients output by the feature pyramid network. The global average feature value is obtained by averaging the feature values ​​at all positions of the feature map.

9. The deep learning method for real-time defect detection in metal additive manufacturing according to claim 1, characterized in that, In step S3, the feature pyramid network generates three feature maps at different scales. After scaling the feature maps at different scales to the same size through bilinear interpolation upsampling, they are spliced ​​together according to the channel dimension to obtain multi-scale fused features. Information entropy analysis is achieved by calculating the probability distribution of feature values ​​of each scale feature map.

10. The deep learning method for real-time defect detection in metal additive manufacturing according to claim 1, characterized in that, In step S5, the anchor frame is set with three scales and three aspect ratios to cover defects of different sizes and shapes. The coordinate transformation is achieved through the calibration matrix of the device coordinate system and the image coordinate system. The calibration matrix is ​​obtained in advance by standard calibration plate.

Citation Information

Patent Citations

  • Additive manufacturing process defect detection method based on multi-sensor data fusion

    CN119334884A

  • PCB mainboard defect detection method and system based on multi-sensor fusion

    CN120655621A