Method for monitoring welding quality of steel stud in whole vehicle manufacturing process based on voiceprint characteristics

By employing a monitoring method based on voiceprint features, combined with PLC signal triggering, sound event detection, and continuous wavelet transform, and using the VGG16 model for welding quality identification, the problem of welding quality monitoring in flexible production lines of automobile manufacturing has been solved, achieving welding quality detection with high accuracy and real-time requirements.

CN120927802APending Publication Date: 2025-11-11ANHUI HUAYUAN INTELLIGENT CONTROL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511029834.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In flexible production lines for automobile manufacturing, the dynamic changes in stud welding points pose significant limitations to image recognition-based monitoring systems. Industrial cameras are difficult to deploy stably, and vibrations cause image blurring. Contact monitoring relying on voltage or current sensors requires modifications to the welding machine circuitry, resulting in high implementation costs and impacting production line cycle time. Existing acoustic signature monitoring methods struggle to separate effective frequency bands under workshop background noise, and EMD decomposition lacks quantitative screening standards, leading to redundant noise in the reconstructed signal and a high false detection rate for deep learning models. Time-frequency analysis based on DWT sacrifices weak high-frequency features, resulting in a high rate of missed detection for non-fusion defects. Complex models such as ResNet have long single-identification times, failing to meet the real-time cycle time requirements of automobile production lines.

Method used

A monitoring method based on acoustic signature is adopted. The audio acquisition system is triggered by PLC signal. Combined with sound event detection technology SED and continuous wavelet transform (CWT), welding quality defects are identified using Moelet wavelet and VGG16 model. This achieves non-contact, non-destructive monitoring without modification. The dual threshold reconstructed signal is screened using permutation entropy (PE) and energy contribution rate (RN). The time-frequency map is optimized to retain the defect feature frequency band by combining Morlet wavelet basis and CWT transform with a scale range of 256. The VGG16 model adopts a two-stage training strategy to meet the real-time beat requirements.

Benefits of technology

It achieves high-accuracy welding quality monitoring in complex workshop environments, solves the monitoring failure problem caused by vibration of moving welding guns, reduces the impact of noise interference, shortens the single identification time, meets the real-time detection needs of automotive production lines, and improves the accuracy and efficiency of welding quality detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120927802A_ABST
    Figure CN120927802A_ABST
Patent Text Reader

Abstract

The invention discloses a method for monitoring the welding quality of a steel stud in the whole vehicle manufacturing process based on voiceprint characteristics. The method comprises the following steps: S1, recording an audio in the welding process in real time, and associating the audio with a vehicle PVI number for storage; s2, positioning and collecting an arc sound key frame in the audio based on a sound event detection technology, and intercepting an audio clip with a preset duration; s3, signal reconstruction processing is carried out on the audio clips which are intercepted for the preset duration; s4, performing continuous wavelet transform on the reconstructed signal to generate a time-frequency graph; s5, inputting the time-frequency graph into a pre-training model for classification, and outputting a welding quality defect identification result; s6, the welding quality defect recognition result is fed back to the PLC production system; according to the method, the problem of dynamic adaptation of welding spots of a production line is solved through simple deployment of the sound pick-up, and electric arc sound positioning is achieved through time sequence modeling; dual thresholds of permutation entropy and energy contribution rate enhance noise immunity, and the feature retention rate is improved; and in combination with a transfer learning strategy, single-point identification is fast, and the comprehensive accuracy is high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of steel stud welding quality technology, specifically a method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics. Background Technology

[0002] In the automotive industry, stud welding is primarily used in the manufacture of body-in-white and other components, supporting the complex and precise assembly requirements of automobile manufacturing. However, non-destructive testing of the welding quality of automated stud welds remains a pressing issue on the production line. Currently, the standard method for detecting the quality of stud welds on the production line is manual disassembly and sampling. In the field of non-destructive testing, the commonly used method is to inspect stud weld quality based on dynamic parameters and weld images; however, this method still has certain shortcomings.

[0003] In current flexible production lines for automobile manufacturing, the dynamic changes in stud welding points pose serious limitations to image recognition-based monitoring systems. Industrial cameras are difficult to deploy stably on moving welding torches, and the image blurring caused by vibration significantly reduces the accuracy of recognition. On the other hand, contact monitoring that relies on voltage or current sensors requires modification of the welding machine circuit, which is costly and affects the production line cycle time. There is an urgent need for a non-contact, modification-free, non-destructive monitoring solution.

[0004] While existing voiceprint detection methods avoid optical limitations, the background noise in the workshop easily drowns out the welding characteristic sound, and traditional Fourier transform is difficult to separate the effective frequency band. Although EMD decomposition has the advantage of adaptability, it lacks a standard for quantitatively screening IMF components, resulting in the reconstructed signal containing redundant noise, which leads to a high false detection rate of deep learning models and cannot meet the high yield requirements of automobile manufacturing.

[0005] Time-frequency analysis based on DWT sacrifices high-frequency weak features to improve speed, resulting in a high rate of missed detection of non-fusion defects; while complex models such as ResNet take a long time to identify each time, far exceeding the cycle time of the body shop, forcing companies to adopt a sampling inspection mode, which is difficult to meet the real-time cycle time requirements of the automotive production line.

[0006] Therefore, it is essential to design a method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics. Summary of the Invention

[0007] The purpose of this invention is to provide a method for monitoring the welding quality of steel studs in the vehicle manufacturing process based on acoustic signature features. This addresses the limitations of image recognition-based monitoring systems in current flexible automotive production lines, where dynamic changes in stud welding points severely restrict image quality. Industrial cameras are difficult to deploy stably on moving welding torches, and vibration-induced image blurring significantly reduces recognition accuracy. Furthermore, contact-based monitoring relying on voltage or current sensors requires modification of the welding machine circuitry, resulting in high implementation costs and impacting production line cycle time. Therefore, a non-contact, modification-free, and non-destructive monitoring solution is urgently needed. While existing acoustic signature monitoring methods avoid optical limitations... However, background noise in the workshop easily drowns out welding characteristic sounds, and traditional Fourier transform is difficult to separate effective frequency bands; although EMD decomposition has adaptive advantages, it lacks a standard for quantitatively screening IMF components, resulting in the reconstructed signal containing redundant noise, leading to a high false detection rate of deep learning models, which cannot meet the high yield requirements of automobile manufacturing. In addition, time-frequency analysis based on DWT sacrifices weak high-frequency features to improve speed, resulting in a high rate of missed detection of non-fusion defects; and complex models such as ResNet take a long time to identify each time, far exceeding the cycle time of the body shop, forcing companies to adopt a sampling inspection mode, which is difficult to meet the real-time cycle time requirements of automobile production lines.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics, comprising the following steps:

[0009] S1: PLC-based welding signal triggered audio acquisition system, which records the welding process audio in real time and stores it in association with the vehicle's PVI number;

[0010] S2: Based on the sound event detection technology SED, key frames of electric arc sound are located and collected in the audio, and electric arc sound audio segments of preset duration are extracted.

[0011] S3: Perform signal reconstruction processing on the arc sound audio segment of preset duration;

[0012] S4: Perform continuous wavelet transform (CWT) on the reconstructed signal to generate a time-frequency graph. The wavelet basis function is the Moelet wavelet, and the scale range is 256.

[0013] S5: Input the time-frequency graph into the pre-trained VGG16 model for classification and output the welding quality defect identification result;

[0014] S6: Feedback the welding quality defect identification results to the PLC production system via the OPC UA protocol.

[0015] As a further technical solution of the present invention, the audio acquisition system in S1 is deployed on the line-side server, and the PLC signal monitoring is implemented through the Spring Boot framework. The audio file naming format is operation time-PVI number.

[0016] As a further technical solution of the present invention, the sound event detection technology SED in S2 specifically includes the following steps:

[0017] S2.1: Extract MFCC coefficients after framing and windowing the original audio segment to extract audio nonlinear features that closely resemble human auditory perception;

[0018] S2.2: The CRNN model is used to train and identify noise features such as electric arc sound and feeder noise. The 40-dimensional MFCC features of the abnormal sound are used as the model input. In the CNN part, 2D convolution is used to process the time-frequency map. In the RNN part, bidirectional LSTM is used to capture the temporal dependence before and after. Finally, a class probability is output at each time step. The feature similarity of each frame is calculated through the CRNN model to locate the start and end points of the electric arc sound.

[0019] S2.3: Extract an arc sound audio segment of preset duration based on the start and end points.

[0020] As a further technical solution of the present invention, the signal reconstruction processing in S3 specifically includes the following steps:

[0021] S3.1: Obtain multiple intrinsic mode functions (IMF) components from the arc sound signal by Empirical Mode Decomposition (EMD).

[0022] S3.2: Calculate the permutation entropy PE and energy contribution rate R of each intrinsic mode function (IMF) component. N ;

[0023] S3.3: Select arrangements that meet the conditions of entropy PE > 0.55 and total energy contribution rate R. N >90% of the intrinsic mode functions (IMF) components are superimposed and reconstructed.

[0024] As a further technical solution of the present invention, in S3.1, Empirical Mode Decomposition (EMD) refers to decomposing a complex signal into several intrinsic mode functions (IMFs) and a residual, thereby achieving signal extraction. As the IMF number increases, the signal gradually loses its original shape and tends to become smoother.

[0025]

[0026] Where I(t) is the input signal; t is time; n is the decomposition level; IMF i (t) represents the intrinsic mode function (IMF) component; r n (t) represents the residual.

[0027] As a further technical solution of the present invention, the entire screening process of the Empirical Mode Decomposition (EMD) mainly involves marking local extreme points, connecting the maximum and minimum points to form upper and lower envelopes, calculating the mean signal lines of the upper and lower envelopes, and subtracting the mean signal lines of the upper and lower envelopes from the input signal to obtain the intermediate signal; after iterating several times, the signals that meet the conditions are obtained as the intrinsic mode functions (IMF) components, and the screening stops when the standard deviation of the following formula is met, thus obtaining the decomposition:

[0028]

[0029] Among them, SD k The standard deviation criterion value represents the result of the k-th selection; t represents the discrete time point; T represents the total length of the signal; h k (t) represents the signal obtained in the k-th screening.

[0030] As a further technical solution of the present invention, the permutation entropy PE and energy contribution rate R in S3.2 are... N The formula is:

[0031] E i =∑|H(IMF) i )| 2

[0032]

[0033] Among them, E i H(IMF) is the energy definition of the i-th IMF. i ) represents performing a Hilbert transform on the i-th IMF; R N This represents the percentage of total energy contributed by the top N IMFs, i.e., the energy contribution rate.

[0034] As a further technical solution of the present invention, the continuous wavelet transform (CWT) formula in S4 is:

[0035]

[0036] Where, ψ * Let ψ(x) be the conjugate complex number; a is the scale factor; b is the time shift factor.

[0037] The Moelet wavelet formula is:

[0038]

[0039] Where t is the time variable; ω0 is the center frequency; when ω≥5, the Morlet wavelet has good bandpass filter characteristics; ω ψ is the standard deviation of the wavelet function, used for normalization to ensure that the energy of the wavelet function is 1.

[0040] As a further technical solution of the present invention, the VGG16 model training in S5 adopts a two-stage strategy:

[0041] The first phase involves freezing the backbone network and training for 50 epochs with a learning rate of 0.001.

[0042] The second phase involves unfreezing the backbone network and training for 50 rounds with a learning rate of 0.0001.

[0043] As a further technical solution of the present invention, the classification of welding quality defect identification results in S5 includes:

[0044] Category 0: No defects;

[0045] Category 1: Burn-through;

[0046] Category 2: Small weld nugget;

[0047] Category 3: Not fused.

[0048] Compared with existing technologies, the beneficial effects of this method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics are:

[0049] To address the challenges of adapting monitoring equipment to the dynamic changes in welding points on the production line, a simple deployment of PLC signal triggering and microphones replaces the traditional industrial camera and sensor modification solution, enabling immediate use in complex car body workshops. The SED technology uses CRNN time-series modeling to dynamically capture the start and end frames of the arc sound, with small positioning errors, solving the monitoring failure problem caused by the vibration of the moving welding gun.

[0050] To address the problem of strong noise in the workshop drowning out effective sound signatures, permutation entropy (PE) and energy contribution rate (R) are used. N The dual-threshold screening of the reconstructed signal improves the feature retention rate compared to the traditional Fourier transform; by combining the Morlet wavelet basis with the CWT transform in the 256-scale range, the required defect feature frequency bands are fully preserved in the time-frequency plot, so that the VGG16 model can still maintain a high accuracy under high background noise interference.

[0051] By combining the VGG16 model training with a two-stage strategy, the time consumed for a single recognition is reduced while ensuring overall accuracy, thus meeting the real-time cycle requirements of the automotive production line. Attached Figure Description

[0052] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0053] Figure 2 This is an unedited image of welding noise.

[0054] Figure 3The intrinsic mode function (IMF) component plots obtained by Empirical Mode Decomposition (EMD).

[0055] Figure 4 A diagram showing the information retention of each intrinsic mode function (IMF) component.

[0056] Figure 5 CWT time-frequency plots of arc characteristics for welding segments with different defects;

[0057] Figure 6 The confusion matrix diagram is based on the optimal model of EMD-CWT-VGG16.

[0058] Figure 7 For training and testing curves;

[0059] Figure 8 This is a schematic diagram of the hardware layout in this invention. Detailed Implementation

[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] Please see the appendix Figure 1 The present invention provides an embodiment of a method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on voiceprint features, comprising the following steps:

[0062] S1: PLC-based welding signal triggered audio acquisition system, which records the welding process audio in real time and stores it in association with the vehicle's PVI number;

[0063] The audio acquisition system is deployed on the lineside server and uses the Spring Boot framework to monitor PLC signals. The audio files are named in the format of job time-PVI number.

[0064] Recorded audio clips such as Figure 2 As shown, specifically, when the PLC's "welding" signal changes to "start", the line-side server recording program starts and continues until the PLC's "welding" signal changes to "end". The program then reads the current vehicle PVI number information from the PLC and stores the file in the format of operation time-PVI number.

[0065] S2: Based on the sound event detection technology SED, key frames of electric arc sound are located and collected in the audio, and electric arc sound audio segments of preset duration are extracted.

[0066] The sound characteristics of stud welding are short and bursting. The audio segment acquired using PLC signals includes the complete trajectory of the welding torch movement from the moment the welding machine receives the welding task to the end of the welding process. Starting at 0.4 seconds, the stud contacts the workpiece, and the welding machine program begins operation. At 0.75 seconds, the arc is ignited; the current and working time at this point are preset by the welding machine to clean surface oil and check circuit quality. At 1 second, the current increases to the set current, a molten pool is formed, and then the stud falls to contact the workpiece. The host machine detects that the voltage difference between the two stages is approximately 0V, and the welding current is cut off. At 1.4 seconds, the welding is complete, and the welding torch is withdrawn from the stud. The audio after 1.5 seconds consists entirely of noise from the welding machine and feeder. Within the 3.5-second segment, only the signal segment from 1 to 1.5 seconds is affected by manually set process parameters, which contains information affecting the stud weld quality. After comparing the acoustic signal characteristics of various steel stud welds, it was found that the duration of the arc acoustic signal of steel stud welds does not exceed 6.5 seconds. In order to completely capture the arc acoustic signal, it was divided according to a ratio of 1:1.2, and the 0.8 seconds before the sound signal is pulled out of the welding torch was selected as the input segment of the defect identification system.

[0067] The preset duration of the arc sound audio segment extracted in the embodiment is set to 800ms;

[0068] Specifically, the following steps are included:

[0069] S2.1: Extract MFCC coefficients after framing and windowing the original audio segment to extract audio nonlinear features that closely resemble human auditory perception;

[0070] S2.2: The CRNN model is used to train and identify noise features such as electric arc sound and feeder noise. The 40-dimensional MFCC features of the abnormal sound are used as the model input. In the CNN part, 2D convolution is used to process the time-frequency map. In the RNN part, bidirectional LSTM is used to capture the temporal dependence before and after. Finally, a class probability is output at each time step. The feature similarity of each frame is calculated through the CRNN model to locate the start and end points of the electric arc sound.

[0071] The following table shows the network structure of the CRNN model:

[0072]

[0073]

[0074] S2.3: Extract an 800ms arc sound audio segment forward from the start and end points;

[0075] S3: Perform signal reconstruction processing on the arc sound audio segment of preset duration;

[0076] Specifically, the following steps are included:

[0077] S3.1: Multiple Intrinsic Mode Function (IMF) components are obtained by decomposing the arc sound signal using Empirical Mode Decomposition (EMD). Empirical Mode Decomposition (EMD) refers to decomposing a complex signal into several IMFs and a residual, thereby achieving signal extraction. As the IMF number increases, the signal gradually loses its original shape and tends to become smoother.

[0078]

[0079] Where I(t) is the input signal; t is time; n is the decomposition level; IMF i (t) represents the intrinsic mode function (IMF) component; r n (t) represents the residual;

[0080] The screening process for Empirical Mode Decomposition (EMD) mainly involves identifying local extrema, connecting the maxima and minima to form upper and lower envelopes, calculating the mean signals of the upper and lower envelopes, and subtracting the mean signals from the input signal to obtain the intermediate signal. This process is iterated several times until the signals satisfying the given conditions are used as IMF components. Screening continues until the standard deviation meets the following equation, at which point the decomposition is obtained.

[0081]

[0082] Among them, SD k The standard deviation criterion value represents the result of the k-th selection; t represents the discrete time point; T represents the total length of the signal; h k (t) represents the signal obtained in the kth screening;

[0083] like Figure 3 As shown, in this embodiment, the arc sound signal is decomposed using Empirical Mode Decomposition (EMD) to obtain 10 Intrinsic Mode Function (IMF) components and 1 residual r. n (t). From IMF1 to IMF10, the signal curve gradually flattens out, and the frequency gradually decreases;

[0084] S3.2: Calculate the permutation entropy PE and energy contribution rate R of each intrinsic mode function (IMF) component. N :

[0085] E i =∑|H(IMF) i )| 2

[0086]

[0087] Among them, E i H(IMF) is the energy definition of the i-th IMF. i) represents performing a Hilbert transform on the i-th IMF; R N This represents the percentage of total energy contributed by the top N IMFs, i.e., the energy contribution rate.

[0088] Permutation entropy (PE) can effectively quantify the local complexity of a signal. Intrinsic Mode Function (IMF) components with high entropy values ​​typically contain rich dynamic information, while IMF components with low entropy values ​​are mainly composed of noise or irrelevant information. The energy contribution rate (R) of each IMF component can be calculated by summing the squares after the Hilbert transform. N The sum of the energies of the intrinsic mode functions (IMF) components;

[0089] S3.3: Select arrangements that meet the conditions of entropy PE > 0.55 and total energy contribution rate R. N >90% of the intrinsic mode functions (IMF) components are superimposed and reconstructed;

[0090] In this embodiment, the permutation entropy PE of each intrinsic mode function (IMF) component is shown in the table below:

[0091]

[0092] like Figure 4 As shown, through permutation entropy (PE) and energy analysis of intrinsic mode functions (IMF) components, it can be found that the permutation entropy (PE) values ​​of the IMF components show a gradually decreasing trend. Among them, IMF1-IMF3 exhibit significantly higher entropy values ​​than other components, with PE>0.55, while the PE values ​​of IMF4 and higher-order components are below 0.5. At the same time, IMF1-3 contribute a total of 94.15% of the energy of the original signal. In addition, the instantaneous frequency range of IMF1-3, 5k-20kHz, covers the characteristic frequency band of welding arc. Therefore, in this embodiment, the signal after IMF4 is discarded, and IMF1 to IMF3 are superimposed to reconstruct the signal and obtain the noise-reduced signal. The new signal removes low-frequency noise but retains the main characteristic information of arc sound, thereby relatively enhancing the characteristic sound located at high frequencies.

[0093] S4: Perform continuous wavelet transform (CWT) on the reconstructed signal to generate a time-frequency graph. The wavelet basis function is the Moelet wavelet, and the scale range is 256.

[0094] The formula for Continuous Wavelet Transform (CWT) is:

[0095]

[0096] Where, ψ * Let ψ(x) be the conjugate complex number; a is the scale factor; b is the time shift factor.

[0097] The Moelet wavelet formula is:

[0098]

[0099] Where t is the time variable; ω0 is the center frequency; when ω≥5, the Morlet wavelet has good bandpass filter characteristics; ω ψ The standard deviation of the wavelet function is used for normalization to ensure that the energy of the wavelet function is 1.

[0100] S5: Input the time-frequency graph into the pre-trained VGG16 model for classification and output the welding quality defect identification results;

[0101] The following table shows the VGG16 model architecture:

[0102]

[0103]

[0104] The VGG16 model training employs a two-stage strategy:

[0105] The first stage freezes the backbone network and trains it for 50 epochs with a learning rate of 0.001; the second stage unfreezes the backbone network and trains it for 50 epochs with a learning rate of 0.0001. The classification of welding quality defect identification results includes:

[0106] Category 0: No defects;

[0107] Category 1: Burn-through;

[0108] Category 2: Small weld nugget;

[0109] Category 3: Incomplete fusion;

[0110]

[0111]

[0112] The table above shows the F1 score of the EMD-CWT-VGG16 model;

[0113] like Figure 6-7As shown, specifically, during the training phase, the input image size is further compressed to a 112×112 color 3-channel RGB image. The pre-trained weights of the backbone network are used for training. First, the model backbone is frozen, and the feature extraction network is not changed; training is performed 50 times with a learning rate of 0.001 and a batch size of 8. Then, the model backbone is unfrozen, and training continues for another 50 times. Due to memory limitations, the learning rate is set to 0.0001, and the batch size is 1. During training, the dataset is divided into a test set and a training set in a 1:9 ratio. 10% of the data in the training set is used as a validation set for verification. Here, 0 represents no defects, 1 represents weld penetration, 2 represents a small weld nugget, and 3 represents no fusion. During the recognition phase, the optimal trained model is selected for classification and recognition.

[0114] S6: Feedback the welding quality defect identification results to the PLC production system via the OPC UA protocol;

[0115] like Figure 8 As shown, in this embodiment, a TAGUAN TSH11 / 0 fully automatic stud welding torch, in conjunction with a TSE11 control system and a TSF11 feeder, is used for automatic welding. No external shielding gas is added throughout the process. Considering the low frequency of weld defects in thicker steel studs, a 0.8mm thick DC54D+Z ultra-deep drawing, stretching, hot-dip galvanized steel plate, which is more prone to defects, is selected as a simulated workpiece. An M6 short-cycle arc welding stud is used for the experiment. The signal acquisition tool is an FHAI NIS-18X omnidirectional microphone with an aluminum alloy housing. It operates in environments from -35℃ to 70℃, has a frequency response range of 20Hz-20kHz (similar to the range of human hearing), a sampling rate of 36kbps / 24bit, and features low latency, small error, strong environmental adaptability, and minimal interference from magnetic fields in the production environment, meeting the experimental requirements. The microphone is connected via UGREEN... The USB 2.0 external sound card connects to the PC. Since the welding torch does not move during single stud welding and the microphone moves with the welding torch on the production site, the microphone is placed without considering the angle with the welding direction, at a distance of 100mm horizontally and 50mm high from the welding point.

[0116] The experimental acquisition program was written using Java and the Spring Boot framework. It set up two interfaces, "Start Recording" and "End Recording", and extracted the audio segment between the two requests, converting the segment into a 16-bit single-channel audio file with a sampling rate of 32000.

[0117] Through experiments, a total of 168 samples were obtained. The causes of these defects include process parameter settings, oil stains on the workpiece surface, and incorrect workpiece placement. In order to address the problem of the small dataset size, the dataset was augmented by random translation and the addition of low-frequency noise, which improved the training accuracy and robustness of the training results.

[0118] The Sound Event Detection (SED) technology is used to locate the key sound coordinates and extract an 800ms segment for subsequent recognition. The SED technology involves dividing the original audio signal into frames, windowing it, and extracting MFCC coefficients. The CRNN model is used to calculate the feature similarity of each frame to find the key frame where the feeder starts to emit noise. An 800ms sound segment is then extracted forward from this frame using coordinates.

[0119] Subsequently, these datasets were uniformly filtered using Empirical Mode Decomposition (EMD). Through permutation entropy (PE) and intrinsic mode function (IMF) component energy analysis, IMF4 and subsequent IMFs were filtered, and IMF1 to IMF3 were superimposed. Then, a two-dimensional time-frequency domain map was generated using Continuous Wavelet Transform (CWT). The wavelet type was selected as Morlet wavelet, the scale range was 256, and the scale values ​​were obtained from a scale array arranged from large to small based on the center frequency and scale range of the wavelet function. After processing all the sample data using EMD and CWT, a self-made arc sound dataset of steel stud weld defects was obtained.

[0120] Because the VGG16 model requires significant computational power to train, to ensure accuracy and efficiency, the input image size was further compressed to a 112×112 color 3-channel RGB image. The model was trained using pre-trained weights from the backbone network. Initially, the backbone was frozen, and the feature extraction network was not modified; training was performed for 50 epochs with a learning rate of 0.001 and a batch size of 8. Then, the backbone was unfrozen, and training continued for another 50 epochs, with the learning rate set to 0.0001 and the batch size to 1. During training, the dataset was divided into a test set and a training set in a 1:9 ratio. 10% of the training set was used as a validation set for verification. The labels were: 0 for no defects, 1 for burn-through, 2 for small weld nugget, and 3 for no fusion. Figure 5 As shown, this is a continuous wavelet transform (CWT) time-frequency diagram of the arc characteristics of welding segments with different defects.

[0121] After training, the deep recognition model can achieve a result of 97.1%. By introducing the confusion matrix and F1-Score as evaluation metrics, the best model trained with VGG16 can find that the main errors occur in the weld penetration defects and small weld nugget defects. The accuracy of all defect recognition can reach 90%, the overall accuracy reaches 96%, and the F1-Score can reach 0.97.

[0122] In summary, this invention addresses the challenge of adapting monitoring equipment to the dynamic changes in welding points on production lines. It replaces the traditional industrial camera and sensor modification scheme by using PLC signal triggering and simple microphone deployment, enabling immediate use in complex car body workshops. The SED technology uses CRNN time-series modeling to dynamically capture the start and end frames of arc sound with small positioning errors, solving the monitoring failure problem caused by the vibration of moving welding guns.

[0123] To address the problem of strong noise in the workshop drowning out effective sound signatures, permutation entropy (PE) and energy contribution rate (R) are used. N The dual-threshold screening of the reconstructed signal improves the feature retention rate compared to the traditional Fourier transform; by combining the Morlet wavelet basis with the CWT transform in the 256-scale range, the required defect feature frequency bands are fully preserved in the time-frequency plot, so that the VGG16 model can still maintain a high accuracy under high background noise interference.

[0124] By combining the VGG16 model training with a two-stage strategy, the time consumed for a single recognition is reduced while ensuring overall accuracy, thus meeting the real-time cycle requirements of the automotive production line.

[0125] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

Claims

1. A method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on voiceprint characteristics, characterized in that: Includes the following steps: S1: PLC-based welding signal triggered audio acquisition system, which records the welding process audio in real time and stores it in association with the vehicle's PVI number; S2: Based on the sound event detection technology SED, key frames of electric arc sound are located and collected in the audio, and electric arc sound audio segments of preset duration are extracted. S3: Perform signal reconstruction processing on the arc sound audio segment of preset duration; S4: Perform continuous wavelet transform (CWT) on the reconstructed signal to generate a time-frequency graph. The wavelet basis function is the Moelet wavelet, and the scale range is 256. S5: Input the time-frequency graph into the pre-trained VGG16 model for classification and output the welding quality defect identification result; S6: Feedback the welding quality defect identification results to the PLC production system via the OPC UA protocol.

2. The method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics according to claim 1, characterized in that: The audio acquisition system in S1 is deployed on the lineside server and uses the Spring Boot framework to monitor PLC signals. The audio file naming format is job time-PVI number.

3. The method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics according to claim 1, characterized in that: The sound event detection technology (SED) in S2 specifically includes the following steps: S2.1: Extract MFCC coefficients after framing and windowing the original audio segment to extract audio nonlinear features that closely resemble human auditory perception; S2.2: The CRNN model is used to train and identify noise features such as electric arc sound and feeder noise. The 40-dimensional MFCC features of the abnormal sound are used as the model input. In the CNN part, 2D convolution is used to process the time-frequency map. In the RNN part, bidirectional LSTM is used to capture the temporal dependence before and after. Finally, a class probability is output at each time step. The feature similarity of each frame is calculated through the CRNN model to locate the start and end points of the electric arc sound. S2.3: Extract an arc sound audio segment of preset duration based on the start and end points.

4. The method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on voiceprint characteristics according to claim 1, characterized in that: The signal reconstruction process in S3 specifically includes the following steps: S3.1: Obtain multiple intrinsic mode functions (IMF) components from the arc sound signal by Empirical Mode Decomposition (EMD). S3.2: Calculate the permutation entropy PE and energy contribution rate R of each intrinsic mode function (IMF) component. N ; S3.3: Select arrangements that meet the conditions of entropy PE > 0.55 and total energy contribution rate R. N >90% of the intrinsic mode functions (IMF) components are superimposed and reconstructed.

5. The method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics according to claim 4, characterized in that: In S3.1, Empirical Mode Decomposition (EMD) refers to decomposing a complex signal into several Intrinsic Mode Functions (IMFs) and a residual, thereby achieving signal extraction. As the IMF number increases, the signal gradually loses its original shape and tends to become smoother. Where I(t) is the input signal; t is time; n is the decomposition level; IMF i (t) represents the intrinsic mode function (IMF) component; r n (t) represents the residual.

6. The method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics according to claim 5, characterized in that: The entire screening process of the Empirical Mode Decomposition (EMD) mainly involves marking local extreme points, connecting the maximum and minimum points to form upper and lower envelopes, calculating the mean signal lines of the upper and lower envelopes, and subtracting the mean signal lines of the upper and lower envelopes from the input signal to obtain the intermediate signal. After iterating several times, the signals that meet the conditions are obtained as the Intrinsic Mode Function (IMF) components. The screening stops when the standard deviation of the following formula is met, thus obtaining the decomposition: Among them, SD k The standard deviation criterion value represents the result of the k-th selection; t represents the discrete time point; T represents the total length of the signal; h k (t) represents the signal obtained in the k-th screening.

7. The method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics according to claim 4, characterized in that: In S3.2, the permutation entropy PE and energy contribution rate R N The formula is: It is i =∑|H(IMF i )| 2 Among them, E i H(IMF) is the energy definition of the i-th IMF. i ) represents performing a Hilbert transform on the i-th IMF; R N This represents the percentage of total energy contributed by the top N IMFs, i.e., the energy contribution rate.

8. The method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics according to claim 1, characterized in that: The continuous wavelet transform (CWT) formula in S4 is as follows: Where, ψ * Let ψ(x) be the conjugate complex number; a is the scale factor; b is the time shift factor. The Moelet wavelet formula is: Where t is the time variable; ω0 is the center frequency; when ω≥5, the Morlet wavelet has good bandpass filter characteristics; ω ψ is the standard deviation of the wavelet function, used for normalization to ensure that the energy of the wavelet function is 1.

9. The method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on acoustic signature characteristics according to claim 1, characterized in that: The VGG16 model training in S5 employs a two-stage strategy: The first phase involves freezing the backbone network and training for 50 epochs with a learning rate of 0.

001. The second phase involves unfreezing the backbone network and training for 50 rounds with a learning rate of 0.0001.

10. The method for monitoring the welding quality of steel studs in the whole vehicle manufacturing process based on voiceprint characteristics according to claim 1, characterized in that: The classification of welding quality defect identification results in S5 includes: Category 0: No defects; Category 1: Burn-through; Category 2: Small weld nugget; Category 3: Not fused.