A Low-Cost Self-Learning Neural Network Design Method for Ultrasonic Detection of Weld Defects
By using an adaptive scaling network and complex domain feature extraction, combined with a self-attention mechanism, the problem of insufficient utilization of complex domain data and high computational cost in existing weld defect detection is solved, achieving efficient and low-cost weld defect identification, which is suitable for intelligent manufacturing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY
- Filing Date
- 2023-03-14
- Publication Date
- 2026-05-26
AI Technical Summary
Existing ultrasonic testing methods for weld defects suffer from problems such as relying on real-domain convolutional neural networks while neglecting complex-domain data, having fixed network architectures, depending on expert experience, and incurring high computational costs, making it difficult to achieve efficient detection of multiple types of defects.
An adaptive scaling network is constructed, which combines aggregation and reduction units with a self-attention mechanism to extract weld defect information using complex domain features. A multi-objective, training-free network search evaluation method is adopted to optimize network performance and reduce computational costs.
It enables efficient identification of weld defects without training, improves detection accuracy, reduces computation time and labor costs, and is suitable for time series prediction and identification tasks, thus promoting the development of intelligent manufacturing.
Smart Images

Figure CN116341621B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of weld defect type identification technology, and more specifically, to a low-cost self-learning neural network design method for ultrasonic detection of weld defects. Background Technology
[0002] Welding defect detection is crucial for construction and industrial activities and is an important task in industrial production. Austenitic stainless steel, as a national resource, makes weld defect detection an even more critical task. Therefore, quickly and timely mastering austenitic stainless steel weld defect detection models is an urgent engineering practice problem to be solved. Ultrasonic testing, due to its high sensitivity, safety, and speed, has become the fastest-growing and most widely used non-destructive testing technology for welds. However, in actual testing, the commonly used methods of manual and machine identification are easily influenced by subjective feelings and experience, and have high computational costs, making them difficult to apply at present. With the advantages of deep learning in image classification, researchers have begun to explore its application in the automated detection of weld defects.
[0003] Currently, deep learning methods proposed by researchers for ultrasonic detection of weld defects can be broadly categorized into two types. One type uses scanned images obtained by the ultrasonic diffraction time-of-flight method as samples for training. In 2019, Huang Huandong et al. analyzed the TOFD-D scan image features of steel plate weld defects and proposed a deep learning model based on a regional convolutional neural network, which automatically identified five types of defects, including weld cracks, porosity, and slag inclusions. In 2022, ZelinZhi et al. proposed a method based on a fast regional deep convolutional neural network, which, through a parallel serial multi-scale feature information fusion mechanism and a channel domain attention strategy, achieved good identification of defects such as cracks and lack of fusion on a titanium alloy welding defect image dataset. While this method has achieved good results, TOFD technology has high requirements for the lateral and longitudinal resolution performance and real-time performance of the ultrasonic probe, making it difficult to apply in practical engineering. Existing mainstream single-sensor ultrasonic A-mode scanning echo detection methods are safer, faster, and have relatively cheaper detection equipment compared to TOFD technology. Therefore, another type of deep learning method for ultrasonic detection of weld defects directly uses ultrasonic A-scan echo signals as samples. In 2020, Silva Lucas C et al. proposed a decision support system based on a deep extreme learning machine, using the spectrum of ultrasonic diffraction time-of-flight signal segments as training features, which can effectively identify four types of weld defects: incomplete fusion, incomplete penetration, slag inclusion, and porosity. In 2021, Thulsiram Gantala et al. introduced an artificial intelligence-based simulation method, automatically creating a dataset based on ultrasonic A-scan echo signals using a small amount of finite element simulation data. The trained convolutional neural network can effectively identify two types of welding defects: porosity and slag inclusion.
[0004] In summary, deep learning based on A-type scan echo samples has been well studied in the field of ultrasonic detection of weld defects, but the following problems still exist: (1) At present, all related studies are carried out on the analysis and extraction of weld defect features on real-domain convolutional neural networks, which ignores the richer complex-domain data forms of the detection samples. Relying only on some real part information for feature extraction and ignoring the nonlinear information on the phase of the imaginary part data and the signal change rate will inevitably affect the integrity and effectiveness of the feature information. The relevant complex convolutional neural network theory and mechanism still need to be improved; (2) The construction and optimization of existing weld detection neural networks are mostly based on adding designed functional modules to the deep backbone network trained in other fields to modify the network architecture to adapt to the defect detection task. This results in a relatively fixed network architecture, which can only be improved by adding, deleting and optimizing the backbone network, which restricts the improvement and expansion of the detection neural network performance; (3) The design of existing deep neural networks for weld detection relies heavily on the subjective experience of experts and generally requires a large amount of data as a training set. It lacks self-supervised learning ability and relies on repeated experiments for structural optimization, which has a high computational cost. The above problems pose technical difficulties for the intelligent and efficient detection of multiple types of weld defects. Therefore, the key technical problem to be solved by this invention is how to explore the collaborative convolution operation mechanism of the real part, real part, and imaginary part of the time-series samples of ultrasonic testing of various weld defects, and to establish a lightweight, high-performance complex convolutional neural network self-learning construction and optimization mechanism. Summary of the Invention
[0005] To address the shortcomings of existing technologies, the present invention aims to provide a low-cost self-learning neural network design method for ultrasonic detection of weld defects. This invention is mainly used to optimize stainless steel weld defect identification technology.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] A low-cost self-learning neural network design method for ultrasonic detection of weld defects is proposed. An adaptive scalable network for weld defect detection is constructed based on weld defect characteristics. A multi-objective, training-free network search and evaluation method is used to evaluate and iteratively search the weld defect detection network to obtain a balanced solution on both network classification accuracy and parameter quantity. Ultimately, this method selects the weld defect type detection network with superior overall performance without any training. The specific steps are as follows:
[0008] Step 1. Acquire, preprocess and select feature domain for weld defect feature signals: Use ultrasonic equipment to acquire the echo one-dimensional ultrasonic signal Xi of weld defects, filter and label it, use the labeled signal to construct a time-domain training dataset and a time-domain test dataset, and map it into the complex domain;
[0009] Step 2. Construct a cell-based adaptive scaling network for weld defect detection: Introduce a self-attention mechanism using aggregation and reduction units to fully extract relevant feature information, further improve the expression of defect feature information, and stack the aggregation and reduction units according to the proposed method to obtain the searched adaptive scaling network;
[0010] Step 3. Use a multi-objective, untrained network search to evaluate the basic network for weld defect detection: Maximize the network classification accuracy and minimize the number of parameters as two competing objectives. At the same time, a low-cost neural network accuracy evaluation strategy is adopted as an evaluation proxy for the network classification accuracy. The dynamic multi-objective balancing optimizer algorithm is used to search and evaluate the adaptive scaling network, so as to select the weld defect type detection network that meets the requirements in terms of comprehensive performance without any training.
[0011] Furthermore, in step 1, the time-domain training dataset and the time-domain test dataset are mapped into the complex domain. The real part is the value of the signal itself, and the imaginary part is the ratio of the negative rate of change to the angular frequency. The real and imaginary parts of the complex domain are used together as the feature vector of weld defects to further characterize the small nonlinear information differences and rates of change in the phase of different types of time series.
[0012] Furthermore, when mapping the time-domain training dataset and the time-domain test dataset into the complex domain, the Morlet complex wavelet transform is used for complex domain mapping. The mapping process is as follows:
[0013] Let the echo be a one-dimensional signal x(t)∈L 2 (R), based on the principle of wavelet transform, the signal x(t) is compared with the wavelet basis. Convolution is performed to obtain W(a,b). The Real and Immagic functions are then used to operate on W(a,b) to obtain the real part matrix Re[W(a,b)] and the imaginary part matrix Im[W(a,b)] in the complex field.
[0014] Furthermore, in step 2, the aggregation unit fuses the defect features in both depth and scale dimensions. Through subsequent aggregation at different levels, the information of the shallow network is refined, enabling jump connections from shallow to deep layers. The specific construction process is as follows:
[0015] Step 2.1 Constructing Aggregation Units: Using jump connections from shallow to deep layers, the candidate nodes of the 6 pending operations are constructed from both depth and scale directions to obtain aggregation cells. Aggregation cells are stacked in N layers to form aggregation units, and downsampling operations are performed between cells.
[0016] Step 2.2 Constructing a candidate network for weld defect detection: The output H(i) of the current aggregate cell is fused with the output H(i-1) of the upper aggregate cell using the correlation feature fusion strategy in the aggregation unit; the aggregation unit and the shrinking unit are stacked and built according to the adaptive scaling network to obtain the candidate network for weld defect detection.
[0017] Step 2.3 Constructing a binary encoding strategy: Encode the eight candidate operations of the aggregation unit, namely 1×1 max pooling, 1×1 convolution, 3×3 average pooling, 3×3 max pooling, 3×3 convolution, 5×5 convolution, 1×3+3×1 convolution, and 1×5+5×1 convolution, using binary notation and setting different strides.
[0018] Step 2.4 Introducing a self-attention mechanism: The real part matrix of the complex field after computation is used as features K, Q, V. The self-attention mechanism is used to enhance the features and synthesize them together with the imaginary part of the complex field to form a single complex matrix, which is used as the complex weld defect depth feature input to the complex fully connected layer for classification.
[0019] Furthermore, in step 2.2, global pooling G is used to obtain the weight vector of the matrix feature information in the current aggregate cell H(i). The weight vector guides the parameter update law of the upper aggregate cell H(i-1). When the number of input data channels is different, the input data is unified through the scale feature unification module SFU. The weld defect features are fused using the following formula:
[0020] F = G(H[i])H[i-1] + H[i].
[0021] Furthermore, in step 2.3, the binary encoding strategy encodes the aggregated cells into binary strings through operations, and uses these strings to specify the decision vector and search space, so as to facilitate subsequent searches for the specific structure of the aggregated units.
[0022] Furthermore, in step 2.3, a self-attention mechanism is used to embed the feature map into a one-dimensional sequence. The length of the features is transformed through a linear layer to ensure that the feature lengths are the same. The calculation formula for the complex fully connected layer is as follows:
[0023]
[0024] In the formula: I represents the number of input neurons in the l-th complex fully connected layer. These represent the input and output of the l-th complex fully connected layer, respectively. and Let represent the weights and biases of the l-th layer, and let i and j represent the positions of the input and output neurons, respectively. Let CReLU represent the complex activation function, with respect to the real part W. r And the imaginary part W iThe complex activation function is calculated as follows:
[0025] CReLU(W) = ReLU(W) r )+iReLU(W i ).
[0026] Furthermore, in step 3, the low-cost neural network accuracy evaluation process is as follows:
[0027] Step 3.1 Network initialization scoring strategy based on feature map features: Input a set of small batches of data of the same category into the searched adaptive scaling candidate network. When initializing the network, extract a feature map for each input image at the same position in the network. When scoring each network, construct the following P difference matrix and calculate the Score. The calculation formula is as follows:
[0028]
[0029] Score = ln|P|
[0030] Where N is the image pixel value;
[0031] Step 3.2 Auxiliary Metrics for Neural Network Accuracy Prediction Based on R-Drop: In the NAS-Bench-201 search space, randomly sample n networks to be evaluated. Calculate the P-difference matrix of the n networks using the P-difference matrix calculation formula under different benchmark image datasets, and obtain the true accuracy of the n networks from the benchmark image datasets. Use the true accuracy as samples to complete the fitting of the R-Drop-based performance predictor. Whenever evaluating the performance of a network, first calculate the network's P-difference matrix and score, and input the P-difference matrix into the fitted R-Drop-based neural network to obtain the prediction accuracy Y of the searched network. Use both the score and Y as indicators to calculate the Grade, determining the performance of the network being evaluated and improving the performance of the evaluation strategy. The Grade calculation formula is as follows:
[0032] Grade = αScore + βY
[0033] Where α and β are weight parameters and α+β=1;
[0034] Step 3.3 Multi-indicator weight comprehensive measurement strategy based on subjective and objective weighting method: Introduce a multi-indicator weight comprehensive measurement strategy based on subjective and objective weighting method to assign weights to the weight parameters in the Grade calculation formula using subjective and objective weighting method.
[0035] Furthermore, the calculation method for the multi-indicator weighted comprehensive measurement strategy is as follows:
[0036] (1) For all x ij Normalize, x ijLet j be the value of the j-th indicator for the i-th sample;
[0037]
[0038] (2) Calculate the objective weight of the i-th sample value under the j-th indicator:
[0039]
[0040] (3) Calculate the entropy value of the j-th index:
[0041]
[0042] (4) Calculate information entropy redundancy:
[0043] d j =1-e j
[0044] (5) Calculate the objective weights of each indicator:
[0045]
[0046] (6) Construct a judgment matrix A, a based on subjective opinions. 1m This indicates the importance of the first indicator relative to the m-th indicator. Each element of the matrix uses numbers 1-9 to represent the relative importance of the two indicators.
[0047]
[0048] (7) Decompose the weights of each indicator based on matrix A and normalize them, calculate the subjective weights of each indicator, and pass the consistency test:
[0049]
[0050] (8) Multi-indicator weighting comprehensive measurement using subjective and objective weighting methods:
[0051]
[0052] Where i = 1, 2, ..., n, j = q = 1, 2, ..., m, n is the number of samples, m is the number of indicators whose weights are to be determined, and k = 1 / ln(n) > 0.
[0053] Furthermore, in step 3, the dynamic multi-objective equilibrium optimizer algorithm searches and evaluates the adaptive scaling network, and optimizes the dynamic multi-objective equilibrium optimizer algorithm by using a ranking strategy based on R2 index contribution, a refraction back learning strategy, and a dynamic exploration and development factor.
[0054] In summary, the invention has the following beneficial effects:
[0055] To verify the effectiveness of the method of this invention, the automatically designed stainless steel weld defect detection network Auto-WDDNet was compared with a manually designed baseline network. The Auto-WDDNet network achieved an accuracy of 96.26% in identifying five types of defects, including slag inclusions, cracks, and porosity, with a single sample detection time of 3.2 ms. It effectively improves defect identification accuracy while reducing the time and manpower costs consumed in network design and training, maintaining a smaller number of parameters and computational consumption. It exhibits better performance than other baseline models, demonstrating the superiority of this method. This invention is applicable to various time-series-based tasks such as prediction, identification, classification, diagnosis, and detection. Furthermore, it helps lower the technical threshold of artificial intelligence, accelerates the engineering application of automated machine learning, and meets the major national needs for promoting the healthy and sustainable development and independent innovation of the intelligent manufacturing industry, thus contributing to the intelligent upgrading of industries and possessing broad market prospects and application value. Attached Figure Description
[0056] Figure 1 This is a flowchart of the present invention;
[0057] Figure 2 This is a schematic diagram of a polymer cell.
[0058] Figure 3 This is a schematic diagram of an aggregation unit;
[0059] Figure 4 This is a schematic diagram of the fusion of related features;
[0060] Figure 5 Flowchart of a multi-objective, untrained network search evaluation method;
[0061] Figure 6 A polka dot plot showing the correlation between the network's true accuracy and the score;
[0062] Figure 7 Diagram of the R-dropout module;
[0063] Figure 8 A scatter plot showing the correlation between the network's true accuracy and Grade.
[0064] Figure 9 The final selected aggregation unit is shown on the left, where the adaptive scaling network framework is on the left and the selected aggregation unit structure is on the right.
[0065] Figure 10 This is a schematic diagram of the self-attention mechanism. Detailed Implementation
[0066] The present invention will now be described in further detail with reference to the accompanying drawings.
[0067] It should be noted that, for ease of description, the descriptions of direction in the following text are consistent with the directions in the accompanying drawings, but they do not limit the structure of the present invention.
[0068] like Figures 1-9 As shown, this invention discloses a low-cost self-learning neural network design method for ultrasonic detection of weld defects. It performs high-dimensional spatial domain feature representation on the initial one-dimensional ultrasonic signal, enriching the feature expression of weld defect data and selecting the feature domain with superior performance. An adaptive scaling network for weld defect detection is constructed based on the weld defect features. From the perspectives of depth and scale, a deep aggregation unit is designed to fuse shallow semantics and deep spatial features, enhancing the network's feature extraction capability. Subsequently, a multi-objective, training-free network search and evaluation method is proposed. This method searches for a large number of aggregation units in the proposed search space. The aggregation units are stacked according to the network stacking method to obtain candidate networks. A low-cost neural network accuracy evaluation strategy is designed to evaluate the network performance. The search and evaluation of candidate networks for weld defect detection are iteratively performed to obtain a balanced solution on both network classification accuracy and parameter quantity. Finally, the method selects the weld defect type detection network with superior overall performance without any training. The specific steps are as follows:
[0069] Step 1. Acquire, preprocess and select feature domains for weld defect feature signals: Use ultrasonic equipment to acquire the echo one-dimensional ultrasonic signal Xi of weld defects, filter and label it, use the labeled signal to construct a time-domain training dataset and a time-domain test dataset, and map them into the complex domain.
[0070] Ultrasonic flaw detectors and ultrasonic angle probes were used to collect and label A-mode echo data from stainless steel weld defect samples (five types: incomplete fusion, porosity, slag inclusion, incomplete penetration, and cracks), constructing time-domain training and testing datasets. To ensure data diversity and validity, the angle probe was kept at a 90-degree angle to the weld centerline during data acquisition, and five A-mode scans were performed: zigzag, forward / backward, left / right, 10°-15° rotation, and circumferential. The probe was moved back and forth within the weld joint cross-section. A Butterworth filter was used to analyze the one-dimensional X-ray signals of the defect detection echoes obtained from the scans. i Filtering was performed to avoid interference from electromechanical coupling on the detection results. The filtered data was then sorted according to the corresponding defect type and stored in five folders. The folders were labeled as cracks, porosity, slag inclusions, lack of fusion, and incomplete penetration, respectively, thus completing the data labeling and constructing a time-domain training dataset. The above steps were repeated to collect data from another batch of stainless steel weld defect samples to construct a time-domain test dataset, which served as the input data for the constructed model.
[0071] One-dimensional weld defect signal samples are mapped into a recursive field (RPF). Detail classes such as points, lines, image colors, or intersection boundaries are used to optimize the original one-dimensional domain, addressing issues like the lack of single defect features and insufficient correlation representation. RPF is an effective method for analyzing the periodicity, chaos, and non-stationarity of signal sequences, and can separate the internal structure, similarity, and predictability of data samples.
[0072] MFCC features have strong noise immunity and excellent performance, and are currently the most effective and widely used features in sequence signal recognition. Therefore, by using a Mel-based array to nonlinearly map the frequency of the original defect signal, and using Mel to replace the original signal frequency, the spectrum is transformed to Mel, which can better reflect the energy distribution characteristics of the weld defect signal. The steps for calculating the Mel cepstral coefficient defect feature matrix include pre-emphasis, windowing and framing, fast Fourier transform (FFT), squaring, Mel filtering, energy superposition, logarithm taking, and discrete cosine transform (DCT). The relationship between the Mel frequency Mel and the weld defect signal frequency is expressed by formula (1), where f represents the frequency.
[0073]
[0074] This invention patent takes into account the applicability of convolutional neural networks (CNNs) to two-dimensional image feature mining, extending the architecture and theoretical foundation of CNNs from the real domain to the complex domain. It combines the powerful feature extraction capabilities of deep CNNs with the comprehensive data representation capabilities of high-dimensional real / complex domain data. This invention maps the one-dimensional ultrasonic detection signal time-domain training dataset and time-domain test dataset into the complex domain, recursive field of effect (RPF), and mellitus-free coefficient-weighted array (MFCC) domain. Based on the detection performance of basic CNNs in a single domain, the invention selects the spatial domain with strong representation capabilities for weld defect information in the corresponding dataset. Experimental results show that the complex domain performs better than the RPF and MFCC domains, and therefore it is used to construct the training and test datasets for the next step, achieving better weld defect detection results.
[0075] The time-domain training and test datasets are mapped into the complex domain. The real part represents the signal value itself, and the imaginary part is the ratio of the negative rate of change to the angular frequency. Introducing the imaginary part information from the complex domain into the weld defect feature vector avoids phase mismatch during information extraction, further characterizes the subtle nonlinear information differences and rates of change in phase of different types of time series, and improves the completeness and effectiveness of defect feature information. The complex domain mapping is performed using Morlet complex wavelet transform, and the mapping process is as follows:
[0076] Let the echo be a one-dimensional signal x(t)∈L 2 (R), based on the principle of wavelet transform, the signal x(t) is compared with the wavelet basis. Convolution is performed to obtain W(a,b). The Real and Imagin functions are then used to operate on W(a,b) to obtain the real part matrix Re[W(a,b)] and the imaginary part matrix Im[W(a,b)] in the complex field, as shown in formulas (2) to (3):
[0077]
[0078]
[0079] In the formula: a is the scale factor, b is the time factor, express .
[0080] Step 2. Construct a cell-based adaptive scaling network for weld defect detection: Introduce a self-attention mechanism using aggregation and reduction units to fully extract relevant feature information, further improve the expression of defect feature information, and stack the aggregation and reduction units according to the proposed method to obtain the searched adaptive scaling network.
[0081] Due to the inherent characteristics of austenitic stainless steel weld materials, ultrasonic testing can cause propagation direction deflection and sound beam attenuation, making defect identification extremely difficult. This results in a low signal-to-noise ratio and a high risk of missed or false defects. To address this, this invention proposes an adaptive weld defect network scaling framework. It designs aggregation units (composed of N layers of stacked aggregate cells) and reduction units, and introduces a self-attention mechanism to further enhance the richness of defect feature information representation. This fully extracts correlated feature information and stacks them according to a predetermined method to obtain a candidate network.
[0082] The aggregation unit includes three levels: such as Figure 2 As shown, the aggregated cell is divided into three levels from top to bottom, fusing defect features in both depth and scale dimensions. Horizontally: Aggregation starts from the shallowest, smallest scale, iteratively absorbing deeper and larger scale features, and refining the information of the shallow network through subsequent aggregation at different levels. Vertically: Skip connections at different depths from shallow to deep layers are designed. Applying the idea of short path reuse, the aggregated cell is constructed into a tree-like connection structure spanning various levels, aggregating all previous feature information (not just the feature information from the previous layer) into subsequent processing. This solves the problem that sequential hierarchical combinations cannot effectively utilize node information, achieving more efficient feature fusion performance than general cascading, while ensuring high accuracy and efficient memory and parameter utilization. The specific construction process is as follows:
[0083] Step 2.1 Constructing Aggregation Units: Using jump connections at different depths from shallow to deep layers, six candidate nodes for alternative operations are constructed from both depth and scale directions to obtain aggregation cells. Aggregation cells are stacked in N layers to form aggregation units, and downsampling operations are performed between cells.
[0084] Step 2.2 Constructing a candidate network for weld defect detection: The aggregation unit contains a correlation feature fusion strategy. The correlation feature fusion strategy in the aggregation unit is used to fuse the output H(i) of the current aggregation cell with the output H(i-1) of the upper aggregation cell. The aggregation unit and the reduction unit are stacked according to the adaptive scaling network to obtain the candidate network for weld defect detection.
[0085] Global pooling G is used to obtain the weight vector of different feature information in the current aggregate cell H(i). The weight vector guides the parameter update law of the upper aggregate cell H(i-1). Weld defect features are fused using the following formula:
[0086] F=G(H[i])H[i-1]+H[i] (4)
[0087] The correlation feature fusion strategy includes a scale feature unification module (SFU) to unify the input data. For example... Figure 3 As shown, since the input data needs to be consistent in both the dimension and channels of the feature map during feature fusion, this invention proposes a Scale Feature Unification Module (SFU) (consisting of an Integer Linear Unit Activation (ReLU) layer, a 1×1 convolutional layer with K kernels, and a batch normalization (BN) layer) to unify the input data when necessary. Here, h, w, and c represent the height, width, and number of channels of the feature map, respectively.
[0088] Step 2.3: To facilitate the search of aggregation units in the subsequent four sections, this invention proposes a binary encoding strategy. Aggregate cells are encoded into binary strings through operations, and these strings are used to specify the decision vector and search space. Detailed information on the operation encoding and the binary substrings are shown in Table 1. Furthermore, the dimensionality of the data in the aggregation unit remains constant, and a "FactorizedReduce" block is used in each reduction unit to halve the feature map size and double the number of channels.
[0089] Table 1. Detailed information on candidate operations
[0090]
[0091]
[0092] Step 2.4 Introducing a self-attention mechanism: The complex field real part matrix after computation is used as the feature matrix K, Q, V, as shown below. Figure 10As shown, a self-attention mechanism is used to enhance features and synthesize a single complex matrix together with the imaginary part of the complex domain. This matrix is then used as the input of the complex weld defect depth feature into the complex fully connected layer for classification. During this process, a relative position encoding mechanism needs to be applied to achieve state reuse and prevent temporal chaos between frames.
[0093] The self-attention mechanism is used to embed the feature map into a one-dimensional sequence. The length of the feature is transformed through a linear layer to make the feature lengths the same. The calculation formula for the complex fully connected layer is as follows:
[0094]
[0095] In the formula: I represents the number of input neurons in the l-th complex fully connected layer. These represent the input and output of the l-th complex fully connected layer, respectively. and Let represent the weights and biases of the l-th layer, and let i and j represent the positions of the input and output neurons, respectively. Let CReLU represent the complex activation function, with respect to the real part W. r And the imaginary part W i The complex activation function is calculated as follows:
[0096] CReLU(W) = ReLU(W) r )+iReLU(W i (6)
[0097] Step 3. Evaluate the base network for weld defect detection using a multi-objective, untrained network search: Traditional weld defect detection network construction uses the accuracy obtained from training the network as the single performance metric. However, the true accuracy of the network requires full training, incurring significant computational costs, and the resulting solution does not adequately balance time and accuracy. To address this issue, this invention proposes a multi-objective, untrained network search and evaluation method. First, a mathematical theory for multi-objective network construction is designed, setting maximizing network classification accuracy and minimizing the number of parameters as two competing objectives. Simultaneously, a low-cost architecture performance evaluation strategy for neural architecture search (LC-NAS) is employed as an evaluation proxy for network classification accuracy. Furthermore, a dynamic multiobjective equilibrium optimizer algorithm (DMOEO) is used to search and evaluate a large number of candidate networks, enabling the selection of the weld defect type detection network with superior overall performance without any training.
[0098] This invention simultaneously considers the network classification accuracy f1(x) and the number of parameters f2(x) as two competing objectives. Maximizing the classification accuracy of the CNN is to improve the network classification performance, while minimizing the parameters is to reduce its computational cost. These two objectives conflict with each other in some cases, that is, the best performance on one objective often leads to a decrease in the performance of the other objective. Under certain constraints, a compromise network between accuracy and computation time is constructed, and the multi-objective formula is as shown in formula (7):
[0099] minF(x)=(f1(x),f2(x)) T (7)
[0100] stx≤θ.
[0101] However, obtaining network classification accuracy requires significant computational resources. Therefore, this invention proposes a low-cost neural network accuracy evaluation strategy (Section 4.2) that has a linear relationship with the test accuracy of the neural network as a proxy for network classification accuracy evaluation. The objective function f1(x) is given by formula (8):
[0102] f1(x) = Grade (8)
[0103] In the formula: Grade is the result of scoring the candidate network by the proposed low-cost neural network accuracy evaluation method, and the specific calculation method is shown in formula (21).
[0104] To eliminate the influence of multiple objectives on the optimization process, the objective function was normalized using formula (9).
[0105]
[0106] In the formula: f i (x) represents the normalized value of the i-th objective function, f i current f represents the value of the i-th objective function during the operation. i min and f i max These represent the minimum and maximum values of the objective function during the operation, respectively.
[0107] The low-cost neural network accuracy evaluation process is as follows:
[0108] Step 3.1 Network Initialization Scoring Strategy Based on Feature Map Features: High-performance networks can better extract feature information unique to different images, resulting in feature maps that differ significantly from the original images. Therefore, for two similar images of the same category with small differences, cm and cn(h) are compared. cm,cn The corresponding feature maps for smaller (H) are also very different. cm,cn(Very large); while poor-performing architectures cannot effectively extract the feature information of similar images, resulting in feature maps that are not much different from or even the same as the original images. Therefore, two similar images of the same category with little difference, cm and cn(h) are very different. cm,cn The corresponding feature maps with smaller differences (H) are also smaller. cm,cn (Smaller). Therefore, H cm,cn -h cm,cn The larger the value, the more image information is extracted by the architecture.
[0109] A set of small batches of data of the same category (such as 128, 64, 32, or 16 similar images) is input into the searched adaptive scaling candidate network. When initializing the network, a feature map is extracted for each input image at the same position in the network. When scoring each network, the following P-difference matrix is constructed and the Score is calculated. The calculation formula is as follows:
[0110]
[0111] Score = ln|P| (11)
[0112] Here, N represents the image pixel value. If a high-performance neural network exists, all elements on the main diagonal will be N, and all off-diagonal elements will be 0 (according to mathematical principles, the Score reaches its maximum when the P-difference matrix is diagonal and all elements on the main diagonal are N). Therefore, high-performance networks have fewer off-diagonal elements. Given two P-difference matrices, the matrix closer to the diagonal will have a higher Score. A higher score during network initialization means higher final accuracy after training. This characteristic can be used to predict the final performance of an untrained network, rather than training the entire network. Figure 6 The scatter plot shows the correlation between the true accuracy and the predicted score of the sampled networks. 1000 networks were sampled from the NAS-Bench-101 benchmark search space and the CIFAR-10 dataset. There is a certain positive correlation between the true accuracy and the score of the network, and there is a certain positive correlation between the true accuracy and the predicted score of the network.
[0113] Step 3.2 Auxiliary metrics for neural network accuracy prediction based on R-Drop:
[0114] The true accuracy and score of a network may exhibit a weak positive correlation and poor search performance due to overfitting of the evaluation strategy. When the strategy is overfitted, the resulting network will perform very well on the training set, but its prediction results on new data will be very unsatisfactory.
[0115] The improved regularization method R-Drop is suitable for the policy evaluation time requirements of this invention. It uses two Dropout iterations and defines a new loss function, Kullback-Leibler (KL) divergence, to control the consistency of predictions, thereby optimizing the network. Specifically, given the input data x for each training step... i , will x i Two forward passes through the network yield two distributions predicted by the network, denoted as P1(y). i |x i ) and P2(y i |x i Because the Dropout operator randomly discards network neurons, the two forward passes are based on two different subnetworks (the same network is used for Dropout, but the missing neurons are different). Therefore, for the same input data (y... i |x i ), P1(y i |x i ) and P2(y i |x i The distributions of the two output distributions of the same sample are different. The R-Drop method then attempts to minimize the two-way KL divergence D between the two output distributions. KL (P1||P2) is used to regularize network prediction, enhancing the network's robustness to Dropout and ensuring consistent output across different Dropout methods. This promotes similarity between "network averaging" and "weight averaging," ultimately improving network performance. The R-Drop network model is as follows: Figure 7 As shown.
[0116] Based on the existing evaluation strategy in step 3.1, an auxiliary indicator for neural network accuracy prediction based on R-Drop is introduced to comprehensively evaluate network performance using multiple indicators. First, n networks are randomly sampled in the NAS-Bench-201 search space. Under different benchmark datasets, the P-difference matrix of the n networks to be evaluated is calculated using formula (10). The true accuracy of the n networks to be evaluated is obtained from NAS-Bench-201 and used as samples to complete the fitting of the R-Drop-based performance predictor. The P-difference matrix constructed in this invention reveals the intrinsic characteristics of all networks. Regardless of the application, the P-difference matrix of superior / inferior networks has the same characteristics. High-performance networks have fewer off-diagonal elements, so the performance predictor can be ported to the corresponding task without retraining. Whenever evaluating the performance of a network, the K-difference matrix and Score of the network are calculated first. The P-difference matrix is then input into the fitted R-Drop-based neural network to obtain the prediction accuracy Y of the network to be evaluated. The Score and Y are used together to calculate the Grade to determine the performance of the network to be evaluated, thereby improving the performance of the evaluation strategy. In this invention, n is 1000, and the Grade calculation formula (12) is shown. α and β are weight parameters and α+β=1.
[0117] Grade = αScore + βY (12)
[0118] Step 3.3 Multi-indicator weight comprehensive measurement strategy based on subjective and objective weighting method: Introduce a multi-indicator weight comprehensive measurement strategy based on subjective and objective weighting method to assign weights to the weight parameters in the Grade calculation formula using subjective and objective weighting method.
[0119] The weight parameters α and β are easily affected by subjective human factors. Based on formula (12), a multi-index weight comprehensive measurement strategy based on subjective and objective weighting method is introduced to assign weights to the weight parameters α and β in formula (12) using subjective and objective weighting method. This takes into account human factors while avoiding deviations caused by human factors to a great extent.
[0120] Subjective weighting methods have a greater advantage than objective weighting methods (entropy weighting) in determining weights based on the decision-maker's intentions, but their objectivity is relatively poor and their subjectivity is relatively strong. Objective weighting methods, on the other hand, have objective advantages, but they cannot reflect the degree of importance that decision-makers attach to different indicators, and they may have certain weights that are contrary to the actual indicators. Addressing the respective advantages and disadvantages of subjective and objective weighting methods, this invention controls subjective randomness within a certain range, achieving an inherent unity of subjective and objective factors, resulting in evaluation results that are truthful, scientific, and credible. Therefore, when assigning weights to indicators, considering the inherent statistical regularities and authoritative values among indicator data, a reasonable decision indicator weighting method is proposed: a multi-indicator weight comprehensive measurement strategy combining subjective and objective weighting methods to compensate for the shortcomings of single weighting methods.
[0121] The calculation method for the multi-indicator weighted comprehensive measurement strategy is as follows:
[0122] (1) For all x ij Normalize, x ij Let j be the value of the j-th indicator for the i-th sample;
[0123]
[0124] (2) Calculate the objective weight of the i-th sample value under the j-th indicator:
[0125]
[0126] (3) Calculate the entropy value of the j-th index:
[0127]
[0128] (4) Calculate information entropy redundancy:
[0129] d j =1-e j (16)
[0130] (5) Calculate the objective weights of each indicator: The weights of each indicator are calculated using the following formula.
[0131]
[0132] (6) Construct a judgment matrix A, a based on subjective opinions. 1m This indicates the importance of the first indicator relative to the m-th indicator. Each element of the matrix uses numbers 1-9 to represent the relative importance of the two indicators.
[0133]
[0134] (7) Decompose the weights of each indicator based on matrix A and normalize them, calculate the subjective weights of each indicator, and pass the consistency test:
[0135]
[0136] (8) Multi-indicator weighting comprehensive measurement using subjective and objective weighting methods:
[0137]
[0138] Where i = 1, 2, ..., n, j = q = 1, 2, ..., m, n is the number of samples, m is the number of indicators whose weights are to be determined, and k = 1 / ln(n) > 0.
[0139] Therefore, formula (12) is updated to formula (21), where w1 and w2 are the subjective and objective weights of the two indicators. Figure 8 The graph shows the correlation between the true accuracy of the sampling network and the predicted score Grade of the multi-index weighting comprehensive measurement strategy based on subjective and objective weighting. The graph shows a strong positive correlation, which proves the effectiveness of the method.
[0140] Grade = w1Score + w2Y (21)
[0141] The dynamic multi-objective equilibrium optimizer algorithm searches and evaluates the adaptive scaling network. It optimizes the algorithm by using a ranking strategy based on R2 contribution, a refraction back learning strategy, and a dynamic exploration and development factor.
[0142] Most current algorithms use crowding distance to sort and delete particles in the non-dominated solution set to update and maintain the external archive. However, this can lead to larger crowding distances and sparser distribution of particles around the deleted particles, making it easier to lose reasonable solutions, and significantly impacting the diversity of the non-dominated solution set. To address these issues, this invention proposes a sorting strategy based on the R² contribution index. Compared with methods such as hypervolume, crowding distance, and TOTSIS, it not only avoids the solution selection pressure caused by Pareto dominance and the problem that crowding distance only represents the distance between individuals and does not reflect the density of individuals, but also effectively reduces the computational workload.
[0143] Assuming the approximate solution set of the Pareto front is A, for ease of calculation, a specific reference point Z is chosen. * The Chebyshev function is used as the utility function, as shown in formula (22). The quality of the solution is evaluated by the contribution value of the R2 index, and the R2 contribution value C of x (x∈A) is used as the utility function. R2 Defined as formula (23).
[0144]
[0145] C R2 (x,A,W,Z * ) = R2(A,W,Z * )-R2(A\{x},W,Z * ) (twenty three)
[0146] In the formula: W is a set of weight vectors in m target spaces, each weight vector w = (w1, w2, ..., w m )∈W are uniformly distributed in the target space. Let W represent the probability distribution over W.
[0147] To further improve the performance of DMOEO, a reverse learning strategy based on convex lens imaging is applied to the archive with a certain probability, causing the current particle to move towards the opposite particle instead of selecting the best particle, in order to avoid getting trapped in local optima. The mathematical model of refraction reverse learning is shown in equation (24).
[0148]
[0149] In the formula: n is the refractive index. The optimal particle is found by adjusting the value of n. In this invention, n is set to 1.2 × 10⁻⁶. 4 UB and LB are the upper and lower bounds of the search space, respectively. i The stored content and particle C i on the contrary.
[0150] In the original EO algorithm, two factors, a1 and a2, dominate the exploration and development process: factor a1 gradually decreases until it disappears at the end of the iteration; factor a2 gradually increases until it reaches its peak at the end of the iteration, thus better balancing the exploration and development processes. However, these values are all constants. Therefore, this invention proposes five iterative equations based on the difference between the exploration and development factors of the current iteration process and the previous generation population during the optimization process. These equations are transformed into dynamic factors to seek a balance between exploration and development during the iteration process, resulting in a better solution. The specific improved equations are shown in Table 2.
[0151] In the formula: c1' and c'2 represent the exploration and development factors of the previous generation population, respectively, to increase the transitivity of intergenerational population correlation factors; a1 and a2 are the factors in the original algorithm. In the improved algorithm, a1 and a2 in the original algorithm are replaced with new factors c1 and c2, respectively.
[0152] Table 2 Improved Equations for Dynamic Exploration and Development Factors c1 and c2
[0153]
[0154] This invention accelerates the application speed of defect reasoning and identification through heterogeneous collaborative computing of "CPU+FPGA".
[0155] To meet the requirements of real-time defect detection and low power consumption, and to accelerate the application speed of weld defect information identification and inference, this invention uses a computing method consisting of computing units with different types of instruction sets and architectures in the network identification and inference stage, namely "CPU+FPGA" heterogeneous collaborative computing. Through high-parallel instruction optimization strategies and high-performance hardware and software co-optimization research, the resource utilization of the FPGA platform is improved, ultimately achieving accelerated inference computing for a lightweight network model for defect detection with high generalization capabilities.
[0156] To achieve an integrated, high-concurrency, and high-performance visualization system for assessment and analysis, a host computer interactive software system and an FPGA hardware acceleration system for stainless steel weld defect detection were designed, developed, and integrated based on application requirements. The host computer software employs a front-end / back-end separation development model, utilizing technologies such as Spring Boot, Redis, Nginx, Spring Cloud, and MySQL to ensure system security, business logic, and data storage. Technologies such as Node.js, Vue.js, and ECharts ensure a visually appealing and smooth system, providing visualized charts of relevant data to improve detection and analysis efficiency, and building a management system for related detection reports and information. Furthermore, the inference accelerator design is based on the Xilinx ZCU104 FPGA evaluation board. The architecture of the Xilinx ZCU104 system is as follows: Figure 5 As shown, this device combines a powerful embedded processing system (PS area) and programmable logic (PL area), featuring 504K system logic units (Look-Up-Table LUTs), 461KCLB flip-flops (FFs), and 38Mb of RAM. It possesses both the flexible and efficient data processing and transaction handling capabilities of an ARM processor and the high-speed parallel processing advantages of an FPGA. During network inference and recognition, a hardware-software co-optimization approach maps the hardware acceleration functions of the PL area to one or more peripherals with specific functions in the PS area. The PS area is primarily responsible for the algorithm scheduling of the entire system and the execution of algorithm modules with complex logical operations, while a large number of repetitive calculations are typically performed by the PL area, utilizing highly parallel instructions to accelerate hardware functions.
[0157] To verify the superiority of DMOEO over current multi-objective optimization algorithms, simulation experiments were conducted to compare it with three baseline algorithms—NSGA-II, NSGA-III, and SMPSO—on ZDT and GLT problems. Optimal values are indicated in bold. Based on the data in Table 3, it is evident that under the same test constraints, DMOEO's statistical results for all 11 test functions are significantly better than the other three comparative algorithms. For test functions F1, F2, F3, F5, F8, and F10, DMOEO consistently obtains the optimal solution on the HV evaluation metric. When solving F2, F3, F5, F8, F9, and F11, DMOEO consistently obtains the optimal solution on the IGD evaluation metric. Although DMOEO failed to find the optimal solution on other test functions, it is still several orders of magnitude better than other algorithms. This demonstrates that the improved DMOEO's pre-Pareto solution is closest to the true Pareto solution compared to other algorithms, exhibiting strong performance and stability. Meanwhile, to avoid potential performance degradation from drastic modifications, detailed ablation experiments were conducted on the modified algorithm to demonstrate its feasibility and correctness. Experimental results show that each modified algorithm exhibits significant performance improvements of varying orders of magnitude in both the HV and IGD evaluation metrics compared to the original algorithm. This proves that the introduced improvement strategies can enhance the algorithm's optimization performance. Therefore, this algorithm can be applied to the candidate network search and evaluation method of this invention and has significant practical value.
[0158] Table 3 Comparison results of different algorithms on the test function
[0159]
[0160]
[0161]
[0162] This invention uses a self-made dataset to compare the proposed LC-NAS method with the NAS-Bench-201 benchmark search space and existing research results. It includes hand-designed methods and common existing NAS methods, and analyzes five types of weld defects. The results are shown in Table 4. All training-free evaluation strategies are combined with random search, with a batch size of 64. The initial population size of the EA method is set to 10, the number of iterations is set to infinity, and the validation accuracy after 12 training epochs is used as the fitness function. The search is paused once the network's search time reaches the expected time (10,000 seconds). Experiments show that the average classification accuracy of the method proposed in this invention reaches 96.22%. Compared with manually designed methods, except for a slight gap compared to SNN methods, the proposed method outperforms other baseline models. However, it is still competitive with SNN methods because neural networks require time-consuming and labor-intensive network design by professionals to achieve the desired results and are highly subjective. The architecture needs to be reconstructed when changing datasets or application contexts. In contrast, the method of this invention, given a baseline search space, can search a large number of networks and output the optimal network in a short time based on the characteristics of different datasets and application contexts without requiring professional design of the neural network. Designing a search space applicable to multiple datasets under the application context and then evaluating the search results will yield even better networks. Furthermore, compared with related NAS methods, although the network accuracy of this invention is slightly lower than that of BlockQNN, its search and evaluation time is several orders of magnitude faster. This verifies the superiority and balanced strategy of the LC-NAS method for weld defect type classification, and it can effectively serve as an evaluation proxy for network classification accuracy.
[0163] Table 4 Comparison of performance with existing studies
[0164]
[0165]
[0166] Figure 9The proposed method is used to select aggregation units when N=3. To fully demonstrate the effectiveness of the proposed method, the proposed Auto-WDDNet network is compared with mainstream manually designed convolutional neural networks SqueezeNet, MobileNetV3, ShuffleNet, DenseNet, ResNet18, ViT, and EfficientNetV2 from different perspectives, including design method, accuracy, number of parameters, and single-sample testing time. The results are shown in Table 5. The proposed method achieves higher classification accuracy than all baseline networks. Compared to the lightweight convolutional neural networks SqueezeNet, MobileNetV3, and ShuffleNet, Auto-WDDNet has approximately half the number of parameters. Compared to conventional convolutional neural networks VGG, DenseNet, and ResNet, Auto-WDDNet has a greater advantage in terms of the number of parameters. Finally, the proposed method, Auto-WDDNet, significantly reduces the number of network parameters while increasing the accuracy by 3.02% compared to the superior EfficientNetV2 network, and the average test time for a single sample is only 3.2ms. This meets the requirements for online real-time identification of stainless steel weld defects and can provide technical support for the development of integrated and portable devices.
[0167] Table 5 Comparison of different network performance experiments
[0168]
[0169] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A low-cost self-learning neural network design method for ultrasonic detection of weld defects, characterized in that: An adaptive scaling network for weld defect detection is constructed based on the characteristics of weld defects. A multi-objective, training-free network search and evaluation method is used to evaluate and iteratively search the weld defect detection network to obtain a balanced solution on both the network classification accuracy and the number of parameters. Finally, the network with the best overall performance for weld defect type detection is selected without any training. The specific steps are as follows: Step 1. Acquire, preprocess and select feature domain for weld defect feature signals: Use ultrasonic equipment to acquire the echo one-dimensional ultrasonic signal Xi of weld defects, filter and label it, use the labeled signal to construct a time-domain training dataset and a time-domain test dataset, and map it into the complex domain; Step 2. Construct a cell-based adaptive scaling network for weld defect detection: A self-attention mechanism is introduced using aggregation and reduction units to fully extract relevant feature information, further improving the representation of defect feature information. The aggregation and reduction units are stacked according to a predetermined method to obtain the searched adaptive scaling network. The aggregation unit fuses defect features in both depth and scale dimensions. Subsequent aggregation at different levels refines the information in the shallow network, achieving jump connections from shallow to deep layers. The specific construction process is as follows: Step 2.1 Constructing Aggregation Units: Using jump connections from shallow to deep layers, the candidate nodes of the 6 pending operations are constructed from both depth and scale directions to obtain aggregation cells. Aggregation cells are stacked in N layers to form aggregation units, and downsampling operations are performed between cells. Step 2.2 Constructing a candidate network for weld defect detection: The output H(i) of the current aggregate cell is fused with the output H(i-1) of the upper aggregate cell using the correlation feature fusion strategy in the aggregation unit; the aggregation unit and the shrinking unit are stacked and built according to the adaptive scaling network to obtain the candidate network for weld defect detection. Step 2.3 Constructing a binary encoding strategy: Encode the eight candidate operations of the aggregation unit, namely 1×1 max pooling, 1×1 convolution, 3×3 average pooling, 3×3 max pooling, 3×3 convolution, 5×5 convolution, 1×3+3×1 convolution, and 1×5+5×1 convolution, using binary notation and setting different strides; Step 2.4 Introduce a self-attention mechanism: use the complex field real part matrix after computation as the feature. K , Q , V The self-attention mechanism is used to enhance features and synthesize a single complex matrix together with the imaginary part of the complex field. This matrix is then used as the input of the complex weld defect depth feature into the complex fully connected layer for classification. Step 3. Use a multi-objective, untrained network search to evaluate the basic network for weld defect detection: Maximizing the network classification accuracy and minimizing the number of parameters are two competing objectives. At the same time, a low-cost neural network accuracy evaluation strategy is used as an evaluation proxy for the network classification accuracy. A dynamic multi-objective balancing optimizer algorithm is used to search and evaluate the adaptive scaling network, so as to select the weld defect type detection network that meets the requirements in terms of comprehensive performance without any training.
2. The low-cost self-learning neural network design method for ultrasonic detection of weld defects according to claim 1, characterized in that: In step 1, the time-domain training dataset and the time-domain test dataset are mapped into the complex domain. The real part is the value of the signal itself, and the imaginary part is the ratio of the negative rate of change to the angular frequency. The real and imaginary parts of the complex domain are used together as the feature vector of weld defects to further characterize the small nonlinear information differences and rates of change in the phase of different types of time series.
3. The low-cost self-learning neural network design method for ultrasonic detection of weld defects according to claim 2, characterized in that, When mapping the time-domain training dataset and time-domain test dataset into the complex domain, the Morlet complex wavelet transform is used for the complex domain mapping. The mapping process is as follows: Assume the echo is a one-dimensional signal Based on the principle of wavelet transform, the signal... With wavelet base Convolution is performed to obtain , respectively functions and function pairs Perform the operation to obtain the real part matrix of the complex field. and imaginary part matrix .
4. The low-cost self-learning neural network design method for ultrasonic detection of weld defects according to claim 1, characterized in that: In step 2.2, global pooling G is used to obtain the weight vector of the matrix feature information in the current aggregate cell H(i). The weight vector guides the parameter update law of the upper aggregate cell H(i-1). When the number of input data channels is different, the input data is unified through the scale feature unification module SFU. The weld defect features are fused using the following formula: 。 5. The low-cost self-learning neural network design method for ultrasonic detection of weld defects according to claim 1, characterized in that: In step 2.3, the binary encoding strategy encodes the aggregated cells into binary strings through operations, and uses these strings to specify the decision vector and search space, so as to facilitate subsequent searches for the specific structure of the aggregated units.
6. The low-cost self-learning neural network design method for ultrasonic detection of weld defects according to claim 1, characterized in that: In step 2.3, a self-attention mechanism is used to embed the feature map into a one-dimensional sequence. The length of the features is transformed through a linear layer to ensure that the feature lengths are the same. The calculation formula for the complex fully connected layer is as follows: In the formula: Indicates the first The number of input neurons in a complex fully connected layer. , They represent the first The inputs and outputs of a complex fully connected layer. and Indicates the first Layer weights and biases, , These represent the positions of the input neuron and the output neuron, respectively. Represents complex activation function For the real part respectively and the virtual part The complex activation function is calculated as follows: 。 7. The low-cost self-learning neural network design method for ultrasonic detection of weld defects according to claim 1, characterized in that: In step 3, the low-cost neural network accuracy evaluation process is as follows: Step 3.1 Network initialization scoring strategy based on feature map features: Input a set of mini-batch data of the same category into the searched adaptive scaling candidate network. When initializing the network, extract a feature map for each input image at the same position in the network. When scoring each network, construct the following... P Difference matrix and calculation Fractions are calculated using the following formula: in, N These are the pixel values of the image. Step 3.2 Auxiliary metrics for neural network accuracy prediction based on R-Drop: Random sampling in the NAS-Bench-201 search space The network to be evaluated passed tests on different benchmark image datasets. P Formula for calculating the difference matrix One network to be evaluated P The difference matrix was obtained from the benchmark image dataset. The true accuracy of each network to be evaluated is used as a sample to complete the performance predictor fitting based on R-Drop. Whenever evaluating the performance of a network, the true accuracy is first calculated... P Difference matrix and and will P The difference matrix is input into a well-fitted R-Drop-based neural network to obtain the prediction accuracy of the searched network. ,use and Calculation of two indicators together To determine the quality of the network under evaluation and improve the performance of the evaluation strategy, the Grade calculation formula is as follows: in, and are weight parameters and ; Step 3.3 Multi-indicator weight comprehensive measurement strategy based on subjective and objective weighting method: The weight parameters in the Grade calculation formula are weighted using the subjective and objective weighting method.
8. The low-cost self-learning neural network design method for ultrasonic detection of weld defects according to claim 7, characterized in that: The calculation method for the multi-index weighted comprehensive measurement strategy is as follows: (1) For all Normalize, For the first The first sample The value of each indicator; (2) Calculate the first The first item under the indicator The objective proportion of each sample value to this indicator: (3) Calculate the first Entropy value of the indicator: (4) Calculate information entropy redundancy: (5) Calculate the objective weights of each indicator: (6) Construct a judgment matrix based on subjective opinions , Indicates the first The first indicator is relative to the first The importance of each indicator is represented by a number from 1 to 9 in the matrix, indicating the relative importance of two indicators. (7) According to the matrix Decompose and normalize the weights of each indicator, calculate the subjective weights of each indicator, and pass the consistency test: (8) Multi-indicator weighting comprehensive measurement using subjective and objective weighting methods: in, , , For the sample size, The number of indicators whose weights need to be determined. .
9. The low-cost self-learning neural network design method for ultrasonic detection of weld defects according to claim 1, characterized in that: In step 3, the dynamic multi-objective balance optimizer algorithm searches and evaluates the adaptive scaling network, and optimizes the dynamic multi-objective balance optimizer algorithm by using a ranking strategy based on R2 index contribution, a refraction back learning strategy, and a dynamic exploration and development factor.