Acoustic emission source positioning method based on lightweight convolution and attention mechanism

Through the improved acoustic emission source localization network (AE-DSCBRNet), combined with depthwise separable convolution and attention mechanism, the accuracy problem of acoustic emission source localization method in high noise environment is solved, high-precision positioning on steel plate structure is achieved, and it has applicability on different materials.

CN120801528AActive Publication Date: 2025-10-17QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)

Patent Information

Application Number
CN202511299393.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-10-17
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing acoustic emission source localization methods are susceptible to interference in high-noise environments, and traditional convolutional neural networks lack effective modeling of complex signal spatiotemporal characteristics, resulting in limited positioning accuracy.

Method used

An improved acoustic emission source localization network (AE-DSCBRNet) is adopted, which combines depth-wise separable convolution blocks, residual convolution blocks and robust channel and spatial attention mechanisms to localize acoustic emission sources through a four-branch convolution structure.

Benefits of technology

High-precision acoustic emission source positioning on steel plate structures is achieved, computational overhead is reduced, and the method is applicable to different materials, thereby improving the robustness and accuracy of positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120801528A_ABST
    Figure CN120801528A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of acoustic emission damage detection, in particular to an acoustic emission source positioning method based on lightweight convolution and an attention mechanism, according to the method, a four-channel continuous wavelet image (CWT) is used as input, a DSConv module carries out integral solution on a standard convolution into channel-by-channel convolution and point-by-point convolution, and the network parameter quantity and the calculation overhead are effectively reduced; the attention module adaptively adjusts feature weights through channel and space attention combination, highlights key time frequency features and inhibits redundant information; the four-branch features integrate multi-channel complementary information through adaptive weighted fusion, so that the signal characterization capability is improved; gradient propagation is improved through residual connection, the deep feature learning efficiency is enhanced, and therefore light-weight and high-precision sound emission source positioning is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of acoustic emission damage detection, and in particular to an acoustic emission source positioning method based on lightweight convolution and attention mechanism. BACKGROUND

[0002] In the fields of power equipment, cable partial discharge monitoring, structural damage detection, etc., traditional positioning methods rely on manual inspection or simple instrument measurement, which is inefficient and difficult to accurately locate the damage source. With the development of acoustic emission technology, damage positioning methods based on acoustic emission signals have become an effective means. This technology captures stress waves generated inside materials due to damage and uses sensor arrays to receive these waves to calculate the coordinates of acoustic emission sources. However, existing acoustic emission source positioning methods have limitations in noise interference, computational complexity, feature extraction, and generalization ability. For example, the time delay estimation (TDOA) based method described in "Acoustic Emission Source Localization Technology" relies on the time difference of signal arrival at different sensors, but is easily disturbed in high noise environments and requires accurate calibration of sensor clock synchronization. Traditional methods usually rely on manual feature extraction, which is difficult to adapt to complex environments, resulting in limited positioning accuracy.

[0003] In recent years, deep learning technology has been increasingly introduced into acoustic emission source positioning tasks. In particular, the successful application of convolutional neural networks (CNN) in image and signal processing provides new possibilities for feature extraction of acoustic emission signals. However, existing methods still have limited feature extraction ability, insufficient noise resistance, and weak generalization ability. For example: traditional CNN uses fixed-size convolution kernels, which is difficult to adapt flexibly to the non-stationary characteristics of acoustic emission signals, and the feature extraction is insufficient; 2D-CNN model (literature: "Single-Sensor Acoustic Emission Source Localization in Plate-Like Structures Using Deep Learning") only uses single-channel signals, which cannot fully utilize multi-sensor information, resulting in the model unable to effectively learn spatial features. The AESLNET model (literature: "Research On The Acoustic Emission Source Localization Methodology In Composite Materials Based On Artificial Intelligence") performs well in composite materials, but this model is mainly trained for specific materials, and its generalization ability in metal materials is weak, resulting in a decrease in positioning accuracy.

[0004] In addition, the physical characteristics of acoustic emission signals have commonalities in different materials: stress waves generated by crack propagation in metal materials and impact waves of delamination defects in composite materials both exhibit high-frequency transient characteristics (frequency range of 100 kHz-1 MHz) and short-time decay rules (millisecond-level duration). This commonality provides theoretical feasibility for deep learning methods based on signal time-frequency characteristics modeling.

[0005] Currently, mainstream acoustic emission source positioning methods include time delay estimation-based methods, pattern matching-based methods, and data-driven deep learning methods. Among them, time delay estimation-based methods rely on the time difference of signals arriving at different sensors, but are easily disturbed in high-noise environments; pattern matching-based methods require the construction of a large number of database matching samples, which has a large amount of calculation; and traditional CNNs lack spatial information modeling capability, so the positioning accuracy is still limited in complex signal environments. SUMMARY

[0006] To address the technical problems of existing acoustic emission source positioning methods in the prior art being easily disturbed in high-noise environments, and traditional convolutional neural networks lacking effective modeling of complex signal spatiotemporal characteristics, resulting in limited positioning accuracy, the present application proposes an acoustic emission source positioning method based on lightweight convolution and attention mechanism. An improved acoustic emission source positioning network (AE-DSCBRNet) is used for accurate positioning of acoustic emission sources in plate structures based on a four-branch convolution structure, combined with deep separable convolution (DSConv) blocks, residual convolution blocks, and robust channel and spatial attention mechanisms.

[0007] The present application is implemented by the following technical solutions: An acoustic emission source positioning method based on lightweight convolution and attention mechanism, comprising the following steps: Step 1: Acoustic emission signals are collected on the sampling grid points on the surface of the steel plate using acoustic emission piezoelectric ceramic sensors; Step 2: Continuous wavelet transform is performed on the acoustic emission signals to convert the acoustic emission time series signals into time-frequency domain image data, and the image is normalized; Step 3: The image data and corresponding acoustic emission source coordinates are integrated into a dataset, which is divided into a training set and a test set, an improved four-branch CNN network (AE-DSCBRNet) model is built and trained; Step 4: The test set is input into the trained network model to obtain the coordinates of the acoustic emission source position, and acoustic emission source positioning is achieved.

[0008] Further, in step 1, a sampling network is divided on the surface of the steel plate, acoustic emission sensors are arranged at the four vertices of the sampling network to form a sensor network, and acoustic emission sources are simulated using a signal generator on the grid points of the sampling to collect acoustic emission signals.

[0009] Further, in step 2, the acoustic emission signal is processed by continuous wavelet transform, wherein the wavelet function is Morse wavelet, and the formula of continuous wavelet transform is:

[0010] wherein, is a continuous complex function in time domain and frequency domain, called mother wavelet; a is a scaling factor, which controls the scaling of the mother wavelet, and b is a shift factor, which controls the position of the wavelet on the time axis.

[0011] Further, in step 2, in order to eliminate the amplitude difference of signals collected by different sensors and improve the numerical stability of CWT time-frequency image, the generated continuous wavelet transform (CWT) time-frequency image is normalized, and the normalization formula is as follows:

[0012] wherein, is the original image, min and max represent the minimum and maximum values of the pixel value respectively, and through the normalization processing, the numerical range of all time-frequency features can be ensured to be consistent, and the influence of uneven data distribution on network training can be avoided.

[0013] Further, the AE-DSCBRNet model includes a data input module, a feature extraction module, a feature fusion module, and an output module. The data input module includes four input branches, each of which receives a continuous wavelet image data. The feature extraction module is composed of two consecutive deep separable convolution (DSConv) blocks, a residual convolution block, and a maximum pooling layer, wherein the DSConv layer extracts local features of the input time-frequency image, the deep convolution captures local patterns in the spatial dimension, and the point-by-point convolution then linearly combines the channel features, thereby significantly reducing the parameter amount and calculation amount while preserving the main time-frequency features of the input image; the residual convolution block further learns the extracted features. The residual convolution block is composed of two consecutive DSConv convolution layers, and the input and output are added through a jump connection to realize the fusion of low-level details and high-level semantics, and batch normalization and ReLU activation function are used after each convolution operation to make the feature distribution more stable and enhance the nonlinear expression ability of the network.

[0014] The mathematical expression of the connection is as follows: ; wherein, x is the input, F(x) is the nonlinear transformation of the convolution layer, and y is the output.

[0015] The target of the residual convolution block is to strengthen deep feature learning while avoiding information decay in deep mapping. After the residual convolution block, the feature map is compressed in space through the max-pooling layer, thereby retaining the main structural information and reducing redundancy; The feature fusion module first splices the feature vectors output by each branch in the feature dimension to form an overall feature representation containing multi-source time-frequency information; then it performs nonlinear mapping on the spliced features through a fully connected layer to learn the importance weights of each branch feature, and generates weighting coefficients using Softmax normalization. Each branch feature vector is multiplied by the corresponding weight and summed to obtain the final comprehensive feature vector; In the output module, the fused comprehensive feature vector is input into two fully connected networks, each of which performs weighted combination and nonlinear mapping on the input feature vector to gradually compress high-dimensional features into low-dimensional representations related to the sound source coordinates. Batch normalization is added after each layer to ensure the consistency of feature distribution in different batches. To avoid gradual decay of information in deep mapping, a residual connection is set after the second fully connected layer to superimpose the initial input image features and the processed features, and the final feature representation is obtained through the ReLU activation function for predicting the acoustic emission source coordinates.

[0016] Further, to enhance the model's attention to spatial information, a CBAM attention mechanism is introduced after the pooling layer, including a channel attention module and a spatial attention module. The channel attention module assigns weights to different channels to highlight features related to the sound source, and the spatial attention module generates an importance distribution of feature regions in the image to guide the network to focus on key time-frequency regions in the image, thereby forming a more discriminative feature representation and improving positioning accuracy.

[0017] Further, the calculation formula of the channel attention is as follows: ; ; The calculation formula of the spatial attention is as follows: ; ; where F is the input feature map, MaxPool(F) and AvgPool(F) represent global average pooling and max pooling, represents the transformation of a 7x7 convolution kernel, represents the Sigmoid function, represents element-wise multiplication, is the final output feature map.

[0018] Further, in the training process of the AE-DSCBRNet model, mean square error (MSE) is used as the loss function, mean absolute error (MAE) is used as the evaluation index, RMSprop is used as the optimizer, and the calculation formula of the mean square error (MSE) is as follows:

[0019]

[0020] wherein n is the sample quantity, is the actual value of the i th sample, is the predicted value of the i th sample; By comparing the difference between the predicted coordinates and the actual coordinates, the mean absolute error (MAE), the root mean square error (RMSE) and the relative error between the two coordinates are used to evaluate the performance of the model.

[0021]

[0022] wherein n is the sample quantity, , is the actual value of the i th sample, , is the predicted value of the i th sample, represents the absolute difference between the actual value and the predicted value.

[0023] The beneficial technical effects of the present application are as follows: (1) The method uses four-channel continuous wavelet image (CWT) as input, the DSConv module decomposes the standard convolution into channel-by-channel convolution and point-by-point convolution, effectively reducing the network parameter quantity and calculation overhead; the attention module adjusts the feature weight through channel and spatial attention, highlights the key time-frequency features and suppresses redundant information; the four-branch feature integrates the multi-channel complementary information through adaptive weighted fusion, improves the signal representation ability; the residual connection improves the gradient propagation and enhances the deep feature learning efficiency, thereby realizing the lightweight and high-precision acoustic emission source positioning.

[0024] (2) The present invention collects acoustic emission signals from the surface of a steel plate and performs continuous wavelet transform processing to convert the time series signals into time-frequency domain image data. At the same time, the image is normalized to obtain a preprocessed time-frequency domain image dataset. On this basis, the present invention constructs an improved convolutional neural network AE-DSCBRNet model, which includes a data input module, a feature extraction module, a feature fusion module, and an output module. Combining lightweight convolution, channel and spatial attention mechanisms, and the design of adaptive multi-branch fusion, the model can obtain more complex features in the acoustic emission signal, thereby achieving precise positioning. Although this experiment was conducted on a steel plate, from the perspective of wave propagation theory, the time-frequency characteristics of the acoustic emission signal have commonalities in different materials. Continuous wavelet transform (CWT) can be used for signal analysis of different materials, and this model can automatically extract key features, so the applicability of this method to different materials has a theoretical basis. Future work will further verify the adaptability of this method in composite materials and heterogeneous structures. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flow chart of the steel plate structure acoustic emission source positioning method described in the present invention, showing the entire process from signal acquisition, preprocessing, model training to positioning prediction; Figure 2 This is the structure diagram of the improved four-branch convolutional neural network described in the present invention, which includes four input branches, a depth-wise separable convolutional feature extraction module, a residual convolution block, a robust channel and spatial attention module, an adaptive multi-branch feature fusion module, and an output module; Figure 3 This is a schematic diagram of the sensor layout and grid on the steel plate surface, showing the layout of four acoustic emission sensors (S1-S4) on the steel plate surface and the sampling grid division; Figure 4 is a schematic diagram of the experimental setup; Figure 5 It is a comparison chart of positioning results, and the deviation range between the predicted value and the actual value is calibrated by the error line. DETAILED DESCRIPTION

[0026] The following will clearly and completely describe the technical solutions in the embodiments of the present invention according to the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0027] like Figure 1 As shown, this embodiment provides a method for locating acoustic emission sources of steel plate structures based on lightweight convolution and attention mechanism, including the following steps: Step 1: Use acoustic emission piezoelectric ceramic sensors to collect acoustic emission signals at sampling grid points on the surface of the steel plate; Specifically, a PZT piezoelectric ceramic sensor (model PXR15, sensitivity >67dB, frequency bandwidth 100kHz-400kHz, resonance frequency 150kHz) is selected as the detection device of the acoustic emission signal. The experimental material is a steel plate with a size of 650mm x 600mm x 3mm, and the surface is divided into a monitoring network of 500mm x 500mm (interval 25mm) Figures 3-4 ). Four acoustic emission sensors are placed at the four corners of the square through a coupling agent, numbered S1, S2, S3, and S4. A coordinate system is established with the center point as the origin, so the coordinates of the sensors are S1(25cm, 25cm), S2(-25cm, 25cm), S3(-25cm, -25cm), and S4(25cm, -25cm).

[0028] Specifically, a signal generator is used to apply a step load to the grid points to simulate real damage source signals. The same number of coordinate points are collected in each quadrant, and to avoid accidental positioning results, each coordinate point is collected multiple times; Step two: continuous wavelet transform is performed on the acoustic emission signal to convert the acoustic emission time series signal into time-frequency domain image data, and the image is normalized; Specifically, the collected acoustic emission signal is processed by continuous wavelet transform, and the formula is:

[0029] where, is the continuous complex function in time and frequency domain, called the mother wavelet. a is the scaling factor, which controls the scaling of the mother wavelet, and b is the shift factor, which controls the position of the wavelet on the time axis.

[0030] Specifically, the image is normalized, and the formula for normalization is:

[0031] Specifically, is the original image, and min and max represent the minimum and maximum values of the pixel value, respectively.

[0032] Step three: integrate the image data and the corresponding acoustic emission source coordinates into a dataset, divide it into a training set and a test set, and build an improved convolutional network with multiple channel inputs for training; Specifically, the continuous wavelet transform (CWT) image is integrated with the corresponding coordinate label to form an acoustic emission signal positioning data set, which contains 960 samples. The data set is divided into a training set and a test set in a ratio of 8:2. In the model training process, the training set and the validation set are input into the multi-layer convolutional network of the application. The network's feature extraction capability and positioning accuracy are optimized by adjusting the training batch, learning rate, and network structure parameters, etc. to obtain a trained model that can accurately predict the acoustic emission source coordinates.

[0033] Specifically, based on the TensorFlow deep learning framework, an improved four-branch CNN network (AE-DSCBRNet) is constructed for high-precision steel plate structure acoustic emission source positioning. The acoustic emission signal is first transformed into a time-frequency image by continuous wavelet transform and input into four parallel branches. The feature extraction part of each branch is composed of a deep separable convolution (DSConv) block, a residual convolution block, and a max-pooling layer. The DSConv extracts local features from the input time-frequency image, the deep convolution captures local patterns in the spatial dimension, and the point-by-point convolution then linearly combines the channel features, thereby significantly reducing the parameter amount and computational amount while preserving the main time-frequency features of the input image. Then, the residual convolution block further learns the extracted features. This module is composed of two consecutive DSConv convolution layers and adds the input and output through a jump connection to realize the fusion of low-level details and high-level semantics. After each convolution operation, batch normalization and ReLU activation function are used to make the feature distribution more stable and enhance the non-linear expression ability of the network.

[0034] The purpose of this module is to strengthen deep feature learning while avoiding information decay in deep mapping. After the residual block, the feature map is spatially compressed by the max-pooling layer to retain the main structural information and reduce redundancy. A robust CBAM attention mechanism is introduced after the pooling layer, including a channel attention module and a spatial attention module. The channel attention module assigns weights to different channels to highlight features related to the acoustic source, and the spatial attention module generates an importance distribution of feature regions in the image to guide the network to focus on key time-frequency regions in the image, thereby forming a more discriminative feature representation and improving the positioning accuracy.

[0035] The calculation formula of the channel attention is as follows: ; ; The calculation formula of the spatial attention is as follows: ; ; where F is the input feature map, MaxPool(F) and AvgPool(F) represent the global average pooling and max pooling, respectively, denotes the transformation of a 7x7 convolution kernel, denotes the Sigmoid function, denotes element-wise multiplication, is the final output feature map.

[0036] Specifically, the four branches respectively perform convolution and pooling processing on the time-frequency images from different sensors to obtain feature vectors representing the local mode and frequency information of the input time-frequency images. Subsequently, these feature vectors are integrated through an adaptive multi-branch fusion module. The module first concatenates the feature vectors output by each branch in the feature dimension to form an overall feature representation containing multi-source time-frequency information. Then, through a fully connected layer, the concatenated features are nonlinearly mapped to learn the importance weights of each branch feature. The weighted coefficients are generated using Softmax normalization. Each branch feature vector is multiplied by the corresponding weight and summed to obtain the final comprehensive feature vector. This module can dynamically adjust the contribution of different sensor inputs and weightedly select and enhance the multi-source time-frequency image features, thereby more effectively responding to key features and improving the accuracy and robustness of acoustic emission source positioning.

[0037] In addition, in the output module, the integrated comprehensive feature vector is input into two layers of fully connected network. Each layer performs weighted combination and nonlinear mapping on the input feature vector, gradually compressing the high-dimensional feature into a low-dimensional representation related to the sound source coordinates. Batch normalization is added after each layer to ensure the consistency of the distribution of image features in different batches. To avoid the gradual attenuation of information in deep mapping, a residual connection is set after the second fully connected layer to superimpose the initial input image features and the processed features, and the final feature representation is obtained through the ReLU activation function for predicting the acoustic emission source coordinates. During model training, the size of the input image is unified to (224, 224, 3), the training batch size is set to 64, the maximum epoch number is set to 900, and the Dropout rate is set to 0.3. The optimizer selects RMSprop, and the loss function uses mean square error (MSE), while the mean absolute error (MAE) is used as the evaluation indicator. In addition, the early stopping strategy (Early Stopping) and the learning rate adaptive adjustment algorithm are used to improve the training efficiency and generalization ability of the model, and the model with the best performance during training is finally saved.

[0038] The structure diagram of the AE-DSCBRNet network is shown in Figure 2 .

[0039] The calculation formulas of the above mean square error MSE and mean absolute error MAE are as follows:

[0040]

[0041] where n is the number of samples, is the actual value of the i-th sample, is the predicted value of the i-th sample.

[0042] Step four: input the test set into the trained network model to obtain the coordinates of the acoustic emission source position, and realize acoustic emission source positioning.

[0043] The performance of the model is evaluated by calculating the error between the predicted coordinates and the actual coordinates, using the mean absolute error (MAE), root mean square error (RMSE), and relative error indicators to quantify the positioning accuracy of the model. The results are shown in Table 1.

[0044]

[0045]

[0046] where n is the number of samples, , is the actual value of the i-th sample, , is the predicted value of the i-th sample, represents the absolute difference between the actual value and the predicted value.

[0047] The predicted coordinates and actual coordinates are shown in Table 1. Figure 5 It can be seen that the distance error between the predicted coordinates and the actual coordinates is less than 12 mm, which is less than half of the 25 mm grid spacing in this study, and this method achieves a resolution of 25 mm.

[0048] As shown in Table 1, Figure 5 this figure shows the difference between the predicted coordinates and the actual coordinates. The results show that the invention can achieve more accurate acoustic emission source positioning.

[0049] Table 1

[0050] To further verify the effectiveness of the present application, comparative experiments were conducted on the same data set with the hyperbolic positioning method based on delay estimation (from Acoustic Emission Source Localization Technology) and different CNN methods (2DCNN model from Single-Sensor Acoustic Emission Source Localization in Plate-Like Structures Using Deep Learning; AESLNET model from Research On The Acoustic Emission Source Localization Methodology In Composite Materials Based On Artificial Intelligence), and the results are shown in Table 2: Table 2

[0051] The experimental results show that the deep separable convolution (DSConv) block, residual convolution block and robust channel and spatial attention mechanism used in the present application can effectively improve the acoustic emission source positioning accuracy. Through the adaptive multi-branch fusion module, the model can fully integrate the complementary features of multi-sensor input and improve the robustness of positioning. Since this method relies on data-driven feature learning rather than specific material physical modeling, it can be extended to acoustic emission source positioning tasks of steel structures, composite materials and other engineering structures after appropriate data migration and fine-tuning, while maintaining lightweight design and low computational overhead.

Claims

1. A method for localizing acoustic emission sources based on lightweight convolution and attention mechanism, characterized by: The steps include: Step 1: Use acoustic emission piezoelectric ceramic sensors to collect acoustic emission signals at sampling grid points on the surface of the steel plate; Step 2: Perform continuous wavelet transform on the acoustic emission signal to convert the acoustic emission time series signal into time-frequency domain image data, and normalize the image; Step 3: Integrate the image data and the corresponding acoustic emission source coordinates into a dataset, divide it into a training set and a test set, build and train the improved four-branch CNN network AE-DSCBRNet model; Step 4: Input the test set into the trained network model to obtain the coordinates of the acoustic emission source and realize the location of the acoustic emission source.

2. The acoustic emission source localization method based on lightweight convolution and attention mechanism according to claim 1, characterized in that: In the step 1, a sampling network is divided on the surface of the steel plate, and acoustic emission sensors are arranged at the four vertices of the sampling network to form a sensor network; a signal generator is used to simulate the acoustic emission source at the divided sampling grid points to collect the acoustic emission signals.

3. The acoustic emission source localization method based on lightweight convolution and attention mechanism according to claim 1, characterized in that: In step 2, the acoustic emission signal is subjected to continuous wavelet transform processing, wherein the wavelet function selects Morse wavelet, and the formula for continuous wavelet transform processing is: ; in, It is a continuous complex function in the time domain and frequency domain, called the mother wavelet; a is the scale factor that controls the scaling of the mother wavelet, and b is the shift factor that controls the position of the wavelet on the time axis.

4. The acoustic emission source localization method based on lightweight convolution and attention mechanism according to claim 1, characterized in that: In step 2, the generated continuous wavelet transform (CWT) time-frequency image is normalized to ensure that the numerical range of all time-frequency features is consistent and to avoid the imbalance of data distribution affecting network training. The normalization formula is as follows: ; in, is the original image, min and max represent the minimum and maximum pixel values ​​respectively.

5. The acoustic emission source localization method based on lightweight convolution and attention mechanism according to claim 1, characterized in that: In step 3, the AE-DSCBRNet model includes a data input module, a feature extraction module, a feature fusion module, and an output module; The data input module contains four input branches, each of which receives a continuous wavelet image data; The feature extraction module consists of two consecutive depth-wise separable convolution DSConv blocks, a residual convolution block, and a maximum pooling layer. The DSConv block extracts local features from the input time-frequency image, while the depth-wise convolution is responsible for capturing local patterns in the spatial dimension. Point-by-point convolution then linearly combines the features of each channel, thereby significantly reducing the number of parameters and computational complexity while retaining the main time-frequency features of the input image. The residual convolution block further learns the extracted features. The residual convolution block consists of two consecutive DSConv convolution layers and adds the input and output through jump connections to achieve the fusion of low-level details and high-level semantics. Each convolution operation is followed by batch normalization and ReLU activation function to make the feature distribution more stable and enhance the nonlinear expression ability of the network. After the residual convolution block, the feature map is spatially compressed through the maximum pooling layer to retain the main structural information and reduce redundancy. The feature fusion module first concatenates the feature vectors output by each branch in the feature dimension to form an overall feature representation containing multi-source time-frequency information. It then performs nonlinear mapping on the concatenated features through a fully connected layer to learn the importance weights of each branch feature. Softmax normalization is then used to generate weighting coefficients. Each branch feature vector is multiplied by the corresponding weight and then summed to obtain the final comprehensive feature vector. In the output module, the fused comprehensive feature vector is input into a two-layer fully connected network. Each layer performs weighted combination and nonlinear mapping on the input feature vector, gradually compressing the high-dimensional features into a low-dimensional representation related to the sound source coordinates. Batch normalization is added after each layer to ensure the consistency of feature distribution of different batches of images. To avoid the gradual attenuation of information in deep mapping, a residual connection is set after the second fully connected layer to superimpose the initial input image features with the processed features, and the ReLU activation function is used to obtain the final feature representation for predicting the coordinates of the acoustic emission source.

6. The acoustic emission source localization method based on lightweight convolution and attention mechanism according to claim 1, characterized in that: In step 3, to enhance the model's ability to focus on spatial information, a CBAM attention mechanism is introduced after the pooling layer, including channel attention and spatial attention modules. The channel attention module is first used to assign weights to different channels to highlight features related to the sound source. The spatial attention module is then used to generate an importance distribution of feature regions in the image, guiding the network to focus on key time-frequency regions in the image, thereby forming a more discriminative feature representation and improving positioning accuracy. The calculation formula of channel attention is as follows: ; ; The calculation formula of spatial attention is as follows: ; ; Among them, F is the input feature map, MaxPool(F) and AvgPool(F) represent global average pooling and maximum pooling respectively. represents the transformation of the 7×7 convolution kernel, represents the Sigmoid function, represents element-wise multiplication, is the final output feature map.

7. The acoustic emission source localization method based on lightweight convolution and attention mechanism according to claim 1, characterized in that: In step 3, during the training of the AE-DSCBRNet model, the mean square error (MSE) is used as the loss function, the mean absolute error (MAE) is used as the evaluation indicator, and RMSprop is used as the optimizer. The calculation formula of the mean square error (MSE) is: ; ; Where n is the number of samples, is the actual value of the i-th sample, is the predicted value of the i-th sample.

8. The acoustic emission source localization method based on lightweight convolution and attention mechanism according to claim 1, characterized in that: In step 4, the difference between the predicted coordinates and the actual coordinates is compared, and the mean absolute error (MAE), root mean square error (RMSE) and the relative error between the two coordinates on the test set are used to evaluate the model performance; ; ; Where n is the number of samples, , is the actual value of the i-th sample, , is the predicted value of the i-th sample, Indicates the absolute difference between the actual value and the predicted value.

Citation Information

Patent Citations

  • Construction method and application of lightweight face mask wearing detection model

    CN114049325A

  • Multi-branch residual convolutional neural network model and image classification method thereof

    CN115205580A

  • Coal rock damage dynamic monitoring and identification method and device

    CN117647586A

  • Hyperspectral image classification method based on hierarchical residual spectrum space convolution network

    CN117893816A

  • Semi-supervised underwater sound event detection method based on joint disturbance consistency constraint

    CN119673210A

Cited By

  • Guided wave signal feature enhancement method based on learnable anti-noise activation

    CN122332938A

  • A guided wave signal feature enhancement method based on learnable anti-noise activation

    CN122332938B