An Augmentation Method for Welding Defect Samples Based on SMOTE_SVM

By optimizing the SMOTE algorithm and SVM hyperparameters, new data samples are generated and optimized support vector machine model is constructed, the problem of unbalanced distribution of welding defect samples is solved and the accuracy of welding defect recognition is improved.

CN115186730BActive Publication Date: 2025-06-20UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210638043.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-07
Publication Date
2025-06-20
Estimated Expiration
2042-06-07

AI Technical Summary

Technical Problem

In the identification of welding defects, since the welding defect samples are far fewer than the weld-forming samples, the samples are unbalanced, and the training effect is not ideal using the support vector machine.

Method used

Using the improved SMOTE algorithm, a new data sample is generated by optimizing the sampling magnification combination of the SMOTE algorithm and the hyperparameters of the SVM, and added it to the original training set to build an optimized support vector machine model to achieve the identification of welding defects.

Benefits of technology

By optimizing the SMOTE algorithm and SVM hyperparameters, the impact of noise samples is reduced, the accuracy of welding defect classification and identification is improved, and effective identification of welding defects is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115186730B_ABST
    Figure CN115186730B_ABST
Patent Text Reader

Abstract

The present invention discloses an augmentation method for welding defect samples based on SMOTE_SVM, belonging to the field of data processing. The sampling magnification combination of the SMOTE algorithm and the hyperparameters of the SVM are optimized to reduce the influence of noise samples during the oversampling process; secondly, the SMOTE algorithm is used to oversample the samples based on the obtained optimal sampling magnification to increase the sample capacity, facilitating the addition of the newly obtained oversampled data samples to the original training set as the new training set in the later stage; finally, model training is carried out based on the SVM with optimized parameters to construct a welding defect recognition model to achieve welding defect recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of unbalanced data classification, and particularly to the field of unbalanced data classification and recognition applied to welding defect recognition. Background Art

[0002] As the carrier device of the new energy vehicle battery pack, the reliability of the welding quality of the battery pack chassis is related to the safety of the battery pack and is closely related to the safe operation of the electric vehicle; at the same time, the battery pack chassis of the new energy vehicle is a typical large-size aluminum alloy part. Due to the characteristics of low melting point and thin thickness of the aluminum alloy part, it is easy to have defects such as spatter and burn-through during processing and welding. Therefore, in order to ensure the reliability of the battery pack, the detection of its welding defects is particularly important. However, with the continuous increase in the demand for automobiles, traditional manual detection faces problems such as too high work intensity and too low efficiency, which puts forward higher requirements for the welding defect detection technology of enterprises. During the welding process of the melting electrode gas shielded welding of aluminum alloy parts, the arc sound signal and the arc electrical signal contain information closely related to welding defects. Therefore, extracting features closely related to welding defects from the sound and electrical signals and constructing a feature set based on this can achieve a comprehensive and diverse description of the welding dynamic process, thereby realizing the recognition of welding defects.

[0003] However, in the actual production process, the available welding defect samples are far less than the samples with good weld formation, which leads to the phenomenon of unbalanced sample distribution. In the case of unbalanced sample distribution, directly using the support vector machine for training has unsatisfactory results. Therefore, it is necessary to perform oversampling on the defect samples to achieve a balanced distribution of the samples. Summary of the Invention

[0004] The Synthetic Minority Oversampling Technique (SMOTE) has been shown in many studies to achieve good results in the problem of imbalanced data classification. However, the SMOTE algorithm cannot selectively choose samples, resulting in being vulnerable to the interference of noisy samples. Therefore, some improved algorithms have been proposed. For example, the Borderline - SMOTE method can selectively choose samples to avoid generating noisy samples. The ASMOTE method avoids the generation of noisy samples by defining a near - neighbor allowable threshold. The above - mentioned studies have improved the blindness of SMOTE. However, like the traditional SMOTE algorithm, all synthesized new samples are fixed multiples of the original samples, unable to achieve precise control of the samples, and also ignoring the influence of SVM hyperparameters in the sampling process. Therefore, this paper simultaneously optimizes the sampling ratio combination of the SMOTE algorithm and the hyperparameters of SVM, and determines the optimal SVM hyperparameters and the new data samples oversampled by the optimal sampling ratio combination under the condition of minimizing the classification error rate. Finally, the new data samples are added to the original training set as the new training set, and the support vector machine with optimized hyperparameters is used for training to establish a welding defect classification and recognition model, ultimately realizing welding defect recognition.

[0005] To achieve the above - mentioned purpose, first, the sampling ratio combination of the SMOTE algorithm and the hyperparameters of SVM are optimized to reduce the influence of noisy samples in the oversampling process; second, the SMOTE algorithm is used to oversample the samples based on the obtained optimal sampling ratio, and the new data samples obtained by oversampling are added to the original training set as the new training set. The technical solution of the present invention is an augmentation method for welding defect samples based on SMOTE_SVM, and this method includes the following steps:

[0006] Step 1: Obtain data samples, classify the sample data set, and establish a support vector machine for each class;

[0007] Step 1.1: Use a Hall sensor to collect welding current signals and use an acoustic sensor to collect arc sound signals during the welding process. To synchronously obtain the above two signals, the sampling frequencies of the acoustic and electrical signals are both set to 48 kHz. After synchronously obtaining the signals, each signal takes 4800 sampling points as a sample. According to the sampling frequency and the number of sampling points of the sample, each sample represents a 0.1 - s welding process. At the same time, according to the welding speed of 10 mm / s, each sample characterizes the welding situation of 1 - mm size. Extract characteristic parameters from the welding current signal samples, and record the obtained welding electrical signal feature vectors as follows:

[0008] A i =[A i,1 ,A i,2 ,…,Ai,m

[0009] where \(i = 1, 2, \ldots, n\), \(n\) is the number of samples; \(m\) is the number of feature parameters; \(A\) i is the \(i\)-th sample of the welding electrical signal. Extract the Mel-frequency cepstral coefficients of the arc sound signal as feature parameters, and denote them as:

[0010] D i = [D i,1 , D i,2 , \ldots, D i,p

[0011] where \(i = 1, 2, \ldots, n\), \(n\) is the total number of frames of the arc sound signal; \(p\) is the dimension of the Mel-frequency cepstral coefficients; \(D\) i is the \(i\)-th sample of the arc sound signal; finally, construct the acoustic-electric sample matrix as follows:

[0012] s i = [A i,1 , \ldots, A i,m×t , D i,1 , \ldots,, D i,p T

[0013] Step 1.2: Construct a multi-class support vector machine, where the number of binary-class support vector machines is determined by the following formula:

[0014]

[0015] where \(k\) is the number of classifications of the dataset;

[0016] Step 2: For each support vector machine, with the minimization of the SVM classification error rate as the optimization objective, the sampling magnification combination of the SMOTE algorithm and the hyperparameters of the SVM as decision variables, construct the objective function as:

[0017] min: \(y = f(X)\), \(X=(C, g, R_1, R_2, \ldots, R\) m )

[0018] In the formula, \(f(X)\) is the classification error rate of the classifier, \(m\) is the number of samples, \(X\) is an individual in the algorithm population, \(R\) i is the sampling magnification of the \(i\)-th sample. \(C\) represents the penalty factor, defined as follows:

[0019] \(C=(C\) max - C\) min )\times r + C\) min , \(i = 1, 2, \ldots, n\)

[0020] \(g\) represents the kernel function parameter, defined as follows:

[0021] \(g=(g\)​​​max -g min )×r + g min , i = 1, 2, …, n

[0022] In the formula, r is a random number with a value range of [0, 1]; C max represents the upper limit of the penalty factor, C min represents the lower limit of the penalty factor; g max represents the upper limit of the kernel function parameter, g min represents the lower limit of the kernel function parameter; The subset of the SMOTE sampling magnification is defined as follows:

[0023]

[0024] where, round(·) is the rounding function, and rand(0, 1) represents a random number between 0 and 1;

[0025] Step 3: Solve the objective function to obtain the optimal SMOTE sampling magnification combination and the optimal SVM hyperparameters;

[0026] Step 3.1: Use Levy shown in the following formula to generate a solution and calculate the objective function value f of the solution i

[0027]

[0028] In the formula, represents the result after the i-th solution is iterated t times; α > 0 is the step size scaling factor; Levy(s, λ) can be expressed as:

[0029]

[0030] where, λ is the Levy flight index, usually taken as 1; The Γ function is a constant for a given λ. For example, when λ is taken as 1, Γ(1 + λ) = 1; The step size s subject to the Levy distribution is expressed according to the Mantegna method as:

[0031]

[0032] In the formula, U ~ N(0, σ 2 ), V ~ N(0, 1), when λ = 1, σ 2 = 1, U ~ N(0, σ 2 ) means that the sample follows a Gaussian normal distribution with a mean of 0 and a variance of σ 2 .

[0033] Step 3.2: Arbitrarily select an objective function solution, and if the fitness f j < f i , then replace the old solution with the new solution, fi denotes the fitness of the \(i\)-th solution;

[0034] Step 3.3: Select the discovered solutions according to the discovery probability, and generate new solutions according to the following formula

[0035]

[0036] where \(s\) is the step size following the Levy distribution, \(H(\cdot)\) is the unit step function, \(p\) a denotes the discovery probability, \(\epsilon\) is a random number drawn from a uniform distribution, are two randomly selected different solutions, denotes the dot product;

[0037] Step 3.4: Save the global optimal solution, and repeat Steps 3.1 to 3.3 until the maximum number of iterations is reached. At this time, the obtained \(X=(C, g, R_1, R_2, \ldots, R\) m ) is the optimal solution obtained by the algorithm, where \(C\) and \(g\) are the optimal penalty factor and kernel parameter of the support vector machine, and \(R_1, R_2, \ldots, R\) m is the optimal sampling magnification combination of SMOTE, where \(m\) is the number of samples.

[0038] Step 4: Based on the optimal sampling magnification of SMOTE obtained in Step 3, use the SMOTE algorithm to oversample the welding defect samples to generate new data samples;

[0039] Step 4.1: For each class sample \(x\), calculate the distance between \(x\) and other samples, and obtain its \(k\) nearest neighbor samples;

[0040] Step 4.2: Let one of the nearest neighbor samples of \(x\) be \(x'\), and synthesize a new sample through the following interpolation method;

[0041] x new =x+(x'-x)×rand(0,1)

[0042] Step 4.3: Add the obtained new data samples to the original training set to synthesize a new training set.

[0043] Furthermore, the characteristic parameters extracted in Step 1.1 include: root mean square value, peak factor, rectified average value, pulse factor, sample variance, power spectrum entropy, peak-to-peak value, energy entropy, waveform factor, singular spectrum entropy, frequency standard deviation.

[0044] In the actual welding production process, the number of welding defect data samples that can be obtained is much less than the number of samples with good weld formation. When the samples are unevenly distributed, directly using a classifier for classification and recognition has an unsatisfactory effect. The traditional oversampling algorithm assigns the same oversampling ratio to all minority class samples, which cannot avoid the generation of noise samples. The present invention proposes a welding defect recognition method based on CS_SMOTE_SVM. This method uses the Cuckoo Search (CS) algorithm to optimize the SMOTE oversampling ratio, giving different sampling ratios to different data samples. This method assigns a lower SMOTE oversampling ratio to noise samples, which can effectively reduce the influence of noise samples during the oversampling process to improve the accuracy of welding defect classification and recognition. Description of the Drawings

[0045] Figure 1 is the flowchart for establishing the welding defect recognition model;

[0046] Figure 2 is the algorithm iteration convergence graph;

[0047] Figure 3 is the classification result graph of 30 experiments;

[0048] Figure 4 is the specific classification result graph of a certain experiment;

[0049] Figure 5 is the schematic diagram of the welding burn-through defect;

[0050] Figure 6 is the schematic diagram of the welding collapse defect. Detailed Embodiments

[0051] The following further illustrates the detailed embodiments of the present invention in conjunction with examples.

[0052] This experiment relies on the welding production process of large-sized thin-walled aluminum alloy cavity parts - aluminum alloy battery box trays. The dataset mainly includes 150 samples with good weld formation, 50 samples with welding spatter defects, and 50 samples with welding burn-through defects. Each sample consists of 4,800 sampling points. According to the sampling rate of 48 kHz, each sample can characterize the welding result of 0.1 s. The welding feature selection method based on CS_SMOTE_SVM involved in the present invention is as Figure 1 shown, and includes the following steps:

[0053] Step 1: Design a class support vector machine.

[0054] Step 1.1: Divide the dataset into three parts, namely the dataset with good weld formation, the welding spatter dataset, and the welding burn-through dataset.

[0055] Step 1.2: Construct a multi-class support vector machine in one-to-one mode, and determine the number of required binary-class support vector machines based on the following formula:

[0056]

[0057] where k is the number of classes, which is taken as 3 in the present invention. Therefore, a total of 3 binary-class support vector machines are required.

[0058] Step 1.2.1: Construct a binary-class support vector machine for classifying and identifying good weld formation and welding spatter. Among them, there are 150 samples of good weld formation and 50 samples of welding spatter.

[0059] Step 1.2.2: Construct a binary-class support vector machine for classifying and identifying good weld formation and welding burn-through. Among them, there are 150 samples of good weld formation and 50 samples of welding burn-through.

[0060] Step 1.2.3: Construct a binary-class support vector machine for classifying and identifying welding spatter and welding burn-through. Among them, there are 50 samples of welding spatter and 50 samples of welding burn-through.

[0061] Step 1.2.4: Take 60% of each type of the above samples as the training set and 40% as the test set.

[0062] Step 2: For each binary support vector machine, with the minimization of the SVM classification error rate as the optimization objective, the sampling magnification combination of the SMOTE algorithm and the hyperparameters of the SVM as the decision variables, construct a parameter optimization model as follows:

[0063] min:y=f(X),X=(C,g,R1,R2,…,R m )

[0064] In the formula, f(X) is the classification error rate of the classifier; m is the number of samples; X is the individual in the algorithm population; R i is the sampling magnification of the i-th sample, and it needs to satisfy minR≤R i ≤maxR, i = 1, 2, …, m; C represents the penalty factor, which is defined as follows:

[0065] C=(C max -C min )×r+C min , i = 1, 2, …, n

[0066] g represents the kernel function parameter, which is defined as follows:

[0067] g=(g max -g min )×r+g min, where \(i = 1, 2, \ldots, n\)

[0068] In the formula, \(r\) is a random number with a value range of \([0, 1]\); \(C\) max represents the upper limit of the penalty factor, and \(C\) min represents the lower limit of the penalty factor; \(g\) max represents the upper limit of the kernel function parameter, and \(g\) min represents the lower limit of the kernel function parameter; the subset of the sampling magnification is defined as follows:

[0069]

[0070] Step 3: Solve the above objective function to obtain the optimal sampling magnification combination and the optimal SVM hyperparameters. First, determine the parameters as shown in the following table:

[0071] Table 1 Algorithm Parameters

[0072]

[0073] Step 3.1: Generate a solution using Levy flight shown in the following formula, and calculate the objective function value \(f\) of the solution i

[0074]

[0075] In the formula, represents the result of the \(i\)-th solution after \(t\) rounds of iteration; \(\alpha>0\) is the step size scaling factor; Levy\((s, \lambda)\) can be expressed as:

[0076]

[0077] The step size \(s\) that follows the Levy distribution can be expressed according to the Mantegna method as:

[0078]

[0079] In the formula, \(U \sim N(0, \sigma\) 2 ), \(V \sim N(0, 1)\), when \(\lambda = 1\), \(\sigma\) 2 = 1.

[0080] Record the \(n\) solution sets randomly generated as where each corresponds to a set of parameters \((C\) i , \(g\) i , \(R1, R2, \ldots, R\) m ). Calculate the objective function values of all and record the optimal solution

[0081] Step 3.2: Arbitrarily select a solution, and the fitness \(f\) of this solution j < fi , then replace the old solution with the new solution. Keep the optimal solution of the previous generation and update the other solutions through Levy(s, λ) to obtain a set of new solutions Compare p k with p k-1 in terms of fitness, and keep the solution with the larger fitness value. At this time, a set of new solutions is obtained

[0082] Step 3.3: Select the discovered solutions according to the discovery probability, and generate new solutions according to the following formula

[0083]

[0084] where α > 0 is the step size scaling factor; s is the step size following the Levy distribution; H(·) is the unit step function; are two randomly selected different solutions; represents the dot product.

[0085] Step 3.4: Save the global optimal solution and repeat the above process until the maximum number of iterations is reached Figure 2 For the comparison of the fitness iteration processes of CS_SMOTE_SVM and GA_SMOTE_SVM, it can be seen from the figure that the fitness curve of the CS_SMOTE_SVM algorithm converges after the 60th iteration, and the optimal fitness value is 85.21%; while the GS_SMOTE_SVM algorithm converges after 53 iterations, and the optimal fitness is 84.3%. This shows that the algorithm can well avoid the local optimal problem, is more conducive to finding the optimal solution in the adaptive search space, and thus enhances the final classification and recognition ability.

[0086] Step 4: Based on the optimal sampling ratio of SMOTE obtained in Step 2, use the SMOTE algorithm to oversample the welding defect samples to generate new data samples;

[0087] Step 4.1: For each sample x, calculate the distance between x and other samples, and obtain its k nearest neighbor samples, where k = 5.

[0088] Step 4.2: Let one of the nearest neighbor samples of x be x', then a new sample can be synthesized through the interpolation method shown in the following formula.

[0089] x new = x + (x' - x) × rand(0, 1)

[0090] Step 4.3: Add the obtained new data samples to the original training set to synthesize a new training set.

[0091] Step 5: Set the hyperparameters of the support vector machine to the optimal hyperparameter values obtained in Step 2, and then train the model based on the new training set to construct a welding defect recognition model.

[0092] Step 6: Pass the test set through the constructed welding defect recognition model to output the classification and recognition results. Figure 3 、 Figure 4 is the experimental result of welding defect recognition, where 3 is the classification accuracy rate of 30 random experiments, and its average classification accuracy rate is 85.21%. Figure 4 is the specific classification situation of each sample in one experiment. Among them, the black square represents the actual type of the sample, and the red circle represents the predicted type of the sample. The higher the coincidence rate of the two curves, the better the classification effect. Among them, the category output results {0, 1, 2} correspond to three types: good weld formation, welding spatter defect, and welding burn-through defect respectively. The confusion matrix is shown in the following table. It can be seen from it that a total of 15 test samples are not correctly classified, and the total classification accuracy rate is 85%. Among them, the classification accuracy rate of good weld formation is 86.7%, the classification accuracy rate of welding spatter defect is 85%, and the classification accuracy rate of welding burn-through defect is 80%. To illustrate the effectiveness of CS_SMOTE_SVM in welding defect recognition, SMOTE_SVM and GA_SMOTE_SVM are used as comparison methods to conduct 30 random experiments under the same conditions. The comparison of G-mean and F-value indicators is shown in Table 3. It can be seen from it that the F-value and G-mean calculated by CS_SMOTE_SVM are better than the other two methods, especially with a relatively high improvement compared to the SMOTE_SVM method. According to the experimental results in Table 3, the F-value of CS_SMOTE_SVM is significantly better than that of SMOTE_SVM, and its G-mean value is also better than the SMOTE_SVM algorithm. This is because the SMOTE_SVM algorithm assigns the same oversampling magnification to each minority class sample, which is likely to generate noise samples far from the decision domain, thus affecting the performance of the learner; the average values of F-value and G-mean of CS_SMOTE_SVM are also better than the GA_SMOTE_SVM algorithm. This is because the GA_SMOTE_SVM algorithm only optimizes the oversampling magnification combination of the SMOTE algorithm, ignoring the role of the penalty factor and kernel parameters in the oversampling process. For different data samples, the optimal hyperparameter values are often different.

[0093] Table 2 Confusion Matrix

[0094]

[0095] Table 3 Comparative Analysis of Welding Defect Recognition Performance

[0096]

Claims

1. An augmentation method for welding defect samples based on SMOTE_SVM, the method comprising the following steps: Step 1: Obtain data samples, classify the sample data set, and establish a support vector machine for each class; Step 1.1: Use a Hall sensor to collect welding current signals and an acoustic sensor to collect arc sound signals during the welding process. To synchronously obtain the above two signals, set the sampling frequencies of the acoustic and electrical signals to 48 kHz. After synchronously obtaining the signals, use 4800 sampling points of each signal as a sample. According to the sampling frequency and the number of sampling points of the sample, each sample represents 0.1 s of the welding process. At the same time, according to the welding speed of 10 mm / s, each sample characterizes the welding situation of 1 mm in size. Extract characteristic parameters from the welding current signal samples, and record the obtained welding electrical signal feature vectors as follows: A i = [A i,1 , A i,2 ,..., A i,m ​ where \(i = 1,2,\cdots,n\), \(n\) is the number of samples; \(m\) is the number of characteristic parameters; \(A\) i is the \(i\)-th sample of the welding electrical signal; extract the Mel-frequency cepstral coefficients of the arc sound signal as characteristic parameters, and denote them as: D i = [D i,1 , D i,2 , …, D i,p ​ where \(i = 1,2,\cdots,n\), \(n\) is the total number of frames of the arc sound signal; \(p\) is the dimension of the Mel frequency cepstral coefficients; \(D\) i is the \(i\)-th sample of the arc sound signal; finally, the acoustic-electric sample matrix is constructed as follows: s i = [A i,1 , …, A i,m×t , D i,1 , …, D i,p T ​ Step 1.2: Construct a multi-class support vector machine, where the number of binary-class support vector machines is determined by the following formula: where k is the number of classifications of the data set; Step 2: For each support vector machine, with the minimization of the SVM classification error rate as the optimization objective, the sampling magnification combination of the SMOTE algorithm and the hyperparameters of the SVM as decision variables, construct the objective function as: min: y = f(X), X = (C, g, R1, R2, …, R m ) where f(X) is the classification error rate of the classifier, m is the number of samples, X is an individual in the algorithm population, and R i is the sampling magnification of the i-th sample. C represents the penalty factor and is defined as follows: C = (C max - C min ) × r + C min , i = 1, 2, …, n g represents the kernel function parameter, defined as follows: g=(g max -g min )×r + g min , i = 1, 2, …, n where r is a random number with a value in the range [0, 1]; C max represents the upper limit of the penalty factor, C min represents the lower limit of the penalty factor; g max represents the upper limit of the kernel function parameter, g min represents the lower limit of the kernel function parameter; The subset of the SMOTE sampling ratio is defined as follows: where round(·) is the rounding function, and rand(0, 1) represents a random number between 0 and 1; Step 3: Solve the objective function to obtain the optimal sampling magnification combination of SMOTE and the optimal sampling magnification of SVM; Step 4: Based on the optimal sampling magnification of SMOTE obtained in Step 3, use the SMOTE algorithm to oversample the welding defect samples to generate new data samples.

2. The augmentation method for welding defect samples based on SMOTE_SVM according to claim 1, characterized in that, The characteristic parameters extracted in Step 1.1 include: root mean square value, peak factor, rectified average value, pulse factor, sample variance, power spectrum entropy, peak-to-peak value, energy entropy, waveform factor, singular spectrum entropy, frequency standard deviation.

3. The augmentation method for welding defect samples based on SMOTE_SVM according to claim 1, characterized in that, The specific method of Step 3 is: Step 3.1: Generate a solution using Levy shown by the following formula, and calculate the objective function value f of the solution i wherein, represents the result after the t-th round of iteration of the i-th solution; α > 0 is the step size scaling factor; Levy(s, λ) is expressed as: where λ is the Levy flight index, taking 1; the Γ function is a constant for a given λ. When λ takes 1, Γ(1 + λ) = 1; The step size s following the Levy distribution is expressed according to the Mantegna method as: wherein, U~N(0,σ 2 ), V~N(0,1), when λ = 1, σ 2 = 1, U~N(0,σ 2 ) means that the sample follows a Gaussian normal distribution with a mean of 0 and a variance of σ 2 ; Step 3.2: Arbitrarily select a solution of the objective function, and the fitness f of this solution j < f i , then replace the old solution with the new solution, where f i represents the fitness of the i-th solution; Step 3.3: Select the discovered solutions according to the discovery probability, and generate new solutions according to the following formula, where s is the step size following the Levy distribution, H(·) is the unit step function, p a represents the discovery probability, ε is a random number drawn from a uniform distribution, are two arbitrarily taken different solutions, represents the dot product; Step 3.4: Save the global optimal solution, and repeat Steps 3.1 to 3.3 until the maximum number of iterations is reached. At this time, the obtained X = (C, g, R1, R2, …, R m ) is the optimal solution obtained by the algorithm, where C and g are the optimal penalty factor and kernel parameter of the support vector machine, and R1, R2, …, R m is the optimal sampling magnification combination of SMOTE, where m is the number of samples.

4. The augmented method for welding defect samples based on SMOTE_SVM according to claim 1, characterized in that, The specific method of Step 4 is: Step 4.1: For each class sample x, calculate the distance between x and other samples, and obtain its k nearest neighbor samples; Step 4.2: Let one of the nearest neighbor samples of x be x', and synthesize a new sample through the following interpolation method; x new = x + (x' - x) × rand(0,1) Step 4.3: Add the obtained new data samples to the original training set to synthesize a new training set.

Citation Information

Patent Citations

  • Ship weld defect detection method based on deep convolutional neural network model

    CN112819806A

  • Method for classifying high-dimensional imbalanced data based on svm

    WO2019041629A1