Acoustic emission nondestructive monitoring method and device, electronic equipment and readable storage medium
By integrating the Transformer and physical information neural network into a joint estimation model, the problems of variable sound velocity and lack of physical constraints in acoustic emission nondestructive monitoring are solved, achieving more accurate sound source localization and medium sound velocity estimation, and improving monitoring accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING UNIV OF CIVIL ENG & ARCHITECTURE
- Filing Date
- 2026-03-24
- Publication Date
- 2026-05-15
AI Technical Summary
Existing acoustic emission nondestructive monitoring methods suffer from significant positioning errors due to the variable sound velocity in non-uniform media. Furthermore, purely data-driven deep learning methods lack physical constraints, resulting in low monitoring accuracy.
A method integrating Transformer and physical information neural network is adopted to jointly estimate the sound source location and the sound velocity in the medium through a joint estimation model. By leveraging the powerful feature learning capability of Transformer and the physical law constraints of physical information neural network, the monitoring accuracy is improved.
It enables more accurate sound source localization and medium sound velocity estimation in non-uniform media, improving the accuracy of non-destructive acoustic emission monitoring.
Smart Images

Figure CN122042826A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of engineering structure condition monitoring technology, and in particular to acoustic emission non-destructive monitoring methods, devices, electronic equipment, and readable storage media. Background Technology
[0002] Acoustic emission (AE), as a dynamic non-destructive testing method, has been widely used in the health monitoring of engineering structures such as concrete. Its core task is to determine the location of internal damage (sound source) through waveform signals received by a sensor array. Existing localization methods are mainly divided into two categories: one is the traditional physical model method based on the time difference of arrival (TDOA) and a preset sound velocity, and the other is the deep learning model method that learns the mapping relationship entirely from the data.
[0003] However, traditional physical model methods rely heavily on the accuracy of preset sound velocity values. In actual non-uniform media, sound velocity is easily variable, and sound velocity mismatch leads to significant positioning errors or even failure. On the other hand, pure data-driven deep learning methods lack physical constraints, and their prediction results often exhibit physical inconsistencies (such as positioning outside the monitoring area). Furthermore, the generalization ability of the model heavily depends on the completeness of the training data, resulting in low accuracy of non-destructive acoustic emission monitoring. Summary of the Invention
[0004] In view of this, the embodiments of this application provide at least a method, apparatus, electronic device and readable storage medium for acoustic emission nondestructive monitoring. By integrating the physical law constraints of Transformer and physical information neural network, the joint estimation of sound source location and medium sound velocity is achieved, thereby improving the accuracy of acoustic emission nondestructive monitoring.
[0005] This application mainly includes the following aspects: In a first aspect, embodiments of this application provide a method for non-destructive acoustic emission monitoring, the method comprising: Acquire the raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored; The original waveform signal is input into the signal preprocessing module of the trained joint estimation model that integrates Transformer and physical information neural network to obtain the preprocessed time domain signal; the joint estimation model also includes a sequentially connected feature extraction module, a fusion encoding module and a multi-task output module; The preprocessed time-domain signal is input into the feature extraction module to obtain the arrival time difference feature vector of the acoustic emission source of the structure to be monitored; The arrival time difference feature vector is input into the fusion coding module to obtain the depth feature representation of the acoustic emission source; The depth feature representation is input to the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated medium equivalent sound velocity of the structure to be monitored.
[0006] Secondly, embodiments of this application also provide an acoustic emission non-destructive testing device, the acoustic emission non-destructive testing device comprising: The data acquisition module is used to acquire the raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored; The data processing module is used to input the original waveform signal into the signal preprocessing module of the trained joint estimation model that integrates Transformer and physical information neural network to obtain the preprocessed time-domain signal; the joint estimation model also includes a sequentially connected feature extraction module, a fusion encoding module and a multi-task output module; The feature extraction module is used to input the preprocessed time-domain signal into the feature extraction module to obtain the arrival time difference feature vector of the acoustic emission source of the structure to be monitored; The feature encoding module is used to input the arrival time difference feature vector into the fusion encoding module to obtain the depth feature representation of the acoustic emission source; The data generation module is used to input the depth feature representation into the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated medium equivalent sound velocity of the structure to be monitored.
[0007] Thirdly, embodiments of this application also provide an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory through the bus, and the machine-readable instructions are executed by the processor to perform the steps of the acoustic emission non-destructive monitoring method as described above.
[0008] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the acoustic emission non-destructive monitoring method as described above.
[0009] The acoustic emission nondestructive monitoring method, apparatus, electronic device, and readable storage medium provided in this application acquire raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored. The raw waveform signals are input into a signal preprocessing module of a trained joint estimation model fusing Transformer and Physical Information Neural Network to obtain a preprocessed time-domain signal. The joint estimation model also includes a sequentially connected feature extraction module, a fusion encoding module, and a multi-task output module. The preprocessed time-domain signal is input into the feature extraction module to obtain the arrival time difference feature vector of the acoustic emission source of the structure to be monitored. The arrival time difference feature vector is input into the fusion encoding module to obtain the depth feature representation of the acoustic emission source. The depth feature representation is input into the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated equivalent sound velocity of the medium in the structure to be monitored. Thus, by fusing the physical constraints of Transformer and Physical Information Neural Network, the joint estimation of the sound source position and the sound velocity of the medium is achieved, improving the accuracy of acoustic emission nondestructive monitoring.
[0010] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A flowchart of a non-destructive acoustic emission monitoring method provided in an embodiment of this application is shown; Figure 2 This illustration shows one of the structural diagrams of a joint estimation model that integrates Transformer and physical information neural network in an embodiment of this application; Figure 3 This is the second schematic diagram of the structure of the joint estimation model that integrates Transformer and physical information neural network in an embodiment of this application; Figure 4 This illustration shows a training diagram of the initial joint estimation model in an embodiment of this application; Figure 5 This illustration shows one of the functional block diagrams of an acoustic emission nondestructive monitoring device provided in an embodiment of this application; Figure 6 This is a second functional block diagram of an acoustic emission non-destructive monitoring device provided in an embodiment of this application; Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0014] To enable those skilled in the art to use the content of this application, and in conjunction with the specific application scenario of "non-destructive monitoring of acoustic emission of concrete materials", the following implementation method is provided. For those skilled in the art, the general principles defined herein can be applied to other embodiments and application scenarios without departing from the spirit and scope of this application.
[0015] The methods, apparatus, electronic devices, or computer-readable storage media described in this application can be applied to any scenario requiring acoustic emission nondestructive testing. This application does not limit specific application scenarios, and any scheme using the acoustic emission nondestructive testing methods and apparatus provided in this application is within the protection scope of this application.
[0016] To facilitate understanding of this application, the technical solutions provided in this application will be described in detail below with reference to specific embodiments.
[0017] Please see Figure 1 , Figure 1 This is a flowchart illustrating a non-destructive acoustic emission monitoring method provided in an embodiment of this application. Figure 1 As shown in the embodiment of this application, the acoustic emission non-destructive monitoring method includes the following steps: S101, acquire the raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored.
[0018] Here, acoustic emission sensors, acting as transducers, are fixed to the surface of the engineering structure to be monitored (such as concrete slabs, beams, columns, etc.), forming a sensor array. When energy is released due to damage initiation or propagation within the structure (e.g., crack formation, fiber breakage), transient elastic waves (i.e., acoustic emission signals) are generated. Each sensor in the array synchronously captures these waves and converts them into a voltage signal that varies with time, i.e., the original waveform signal.
[0019] In this embodiment, a 300mm × 300mm × 30mm concrete slab specimen is used as the structure to be monitored. Four acoustic emission sensors (model RS-2A) are arranged on its surface, forming a rectangular array. Acoustic emission events are simulated at different locations on the specimen surface through a lead-breaking test (acoustic emission instrument model DS5), and the data acquisition system synchronously records the waveform of each sensor channel. A typical acoustic emission event will generate raw waveform data containing four channels. These waveforms differ in arrival time, amplitude, and shape, and these differences contain information about the sound source location and medium properties.
[0020] S102, the original waveform signal is input to the signal preprocessing module of the trained joint estimation model of the fusion Transformer and physical information neural network to obtain the preprocessed time domain signal; the joint estimation model also includes a sequentially connected feature extraction module, fusion encoding module and multi-task output module.
[0021] Here, since the raw waveform signals acquired on-site are usually not directly usable for high-precision analysis, this method uses a pre-trained dedicated neural network model for processing. Please refer to [link / reference needed]. Figure 2 , Figure 2 This is one of the schematic diagrams of the joint estimation model integrating Transformer and physical information neural network in the embodiments of this application. Figure 2 As shown, the joint estimation model 200 includes a sequentially connected signal preprocessing module 210, a feature extraction module 220, a fusion encoding module 230, and a multi-task output module 240. Its core innovation lies in integrating the powerful feature learning capabilities of the Transformer architecture with the physical constraints of a Physical Information Neural Network (PINN). The raw waveform first enters the signal preprocessing module 210, which automatically cleans and normalizes the input signal to improve signal quality and provide standardized input for subsequent core calculations.
[0022] In this embodiment, the joint estimation model is trained offline. Its signal preprocessing module can automatically process the input multi-channel raw waveforms and output a set of time-domain signals with uniform format and improved quality. This step lays the foundation for subsequent feature extraction and joint estimation.
[0023] S103, the preprocessed time-domain signal is input to the feature extraction module to obtain the arrival time difference feature vector of the acoustic emission source of the structure to be monitored.
[0024] Here, the preprocessed signal is fed into the feature extraction module 220 of the model. The core function of this module is to automatically extract a key feature from the standardized waveforms of each channel that can characterize the spatial geometric relationship between the sound source and the sensor array—that is, the arrival time difference of the sound waves to different sensors. This feature vector is the key input for subsequent sound source localization and state inversion.
[0025] In this embodiment, after processing the preprocessed waveforms of the four channels, the feature extraction module 220 outputs a multidimensional feature vector. This vector is essentially a compressed mathematical representation of the continuous waveform data, characterizing the spatiotemporal relationship of sound wave propagation in the infrasonic emission event, for use by the subsequent complex fusion coding module 230.
[0026] S104, the arrival time difference feature vector is input to the fusion coding module to obtain the depth feature representation of the acoustic emission source.
[0027] Here, the time difference of arrival feature vector is then input into the core of the model—the fusion encoding module 230. This module utilizes mechanisms such as self-attention in the Transformer architecture to deeply mine and fuse the complex inter-sensor dependencies and global information contained in the feature vector. Through this process, the simple time difference feature is transformed into a highly abstract and information-rich deep feature representation, which fuses all the high-order information needed to infer the sound source location and the sound velocity in the medium.
[0028] In this embodiment, the fusion coding module 230 generates a novel deep feature representation by performing deep nonlinear transformation and information fusion on the time difference of arrival features. This step is crucial for elevating traditional physical features to a high-level representation capable of supporting complex joint estimation tasks.
[0029] S105, the depth feature representation is input to the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated medium equivalent sound velocity of the structure to be monitored.
[0030] Here, the depth feature representation is ultimately passed to the model's multi-task output module 240. This module is designed to solve a joint estimation problem: based on the same depth feature, it simultaneously outputs two physically meaningful quantities. The first output is an estimate of the two-dimensional position coordinates of the acoustic emission event within the monitoring plane, directly indicating the spatial location of the damage. The second output is an estimate of the equivalent sound velocity of the medium under the current monitoring environment, an important physical parameter reflecting the internal state of the material.
[0031] In this embodiment, the multi-task output module 240 outputs two results at once. Compared with traditional techniques, this method based on a fusion model and joint estimation framework can achieve more accurate sound source localization and simultaneously obtain valuable medium sound velocity information, thereby achieving a deeper understanding of the structural state while completing the localization.
[0032] Further, the signal preprocessing module includes a bandpass filtering submodule, a signal normalization submodule, and a length normalization submodule; the signal preprocessing module that inputs the original waveform signal into the trained joint estimation model fusing Transformer and physical information neural network to obtain the preprocessed time-domain signal includes: Step a1: Input the original waveform signal into the bandpass filter submodule to obtain the effective frequency band signal after filtering out environmental noise.
[0033] Please see here. Figure 3 , Figure 3 This is the second schematic diagram of the joint estimation model integrating Transformer and physical information neural network in the embodiments of this application. Figure 3 As shown, the signal preprocessing module 210 includes a bandpass filtering submodule 211, a signal normalization submodule 212, and a length normalization submodule 213. Bandpass filtering is the first step in signal preprocessing. Its purpose is to selectively retain effective frequency bands related to the acoustic emission characteristics of the material from the original signal containing multiple frequency components, while suppressing irrelevant or interfering components such as low-frequency mechanical vibrations and high-frequency electronic noise. This is equivalent to creating a clean spectral window for subsequent analysis.
[0034] In this embodiment, considering the main energy distribution of acoustic emission signals in concrete materials, the bandpass filter submodule 211 is configured as a Butterworth filter with a passband frequency range of [50kHz, 400kHz]. This filter processes the acquired original waveform signals of each channel, filtering out frequency components outside the passband, thereby obtaining a frequency band signal containing only effective acoustic emission information.
[0035] Step a2: Input the effective frequency band signal into the signal normalization submodule to obtain the amplitude-normalized signal.
[0036] Here, the purpose of signal normalization is to eliminate or reduce the inconsistency in signal amplitude caused by factors such as differences in sensor sensitivity, different signal propagation path attenuation, and preamplifier gain settings. By scaling the amplitude, all signals are placed within a uniform amplitude scale, avoiding bias in the model due to differences in amplitude dimensions, and focusing on the extraction of waveform morphology and timing features.
[0037] In this embodiment, the signal normalization submodule 212 performs a maximum-minimum normalization operation on the filtered signal of each channel. Specifically, it calculates the absolute maximum value of the channel signal within an event time window, and then divides all data points of the channel by the maximum value, so that the amplitude range of the signal is linearly scaled to the interval [-1, 1], thereby obtaining the amplitude-normalized signal.
[0038] Step a3: Input the amplitude-normalized signal into the length-normalization submodule to obtain the preprocessed time-domain signal with a uniform time length.
[0039] Here, the purpose of length normalization is to ensure that all sample data input to subsequent neural network models have the same dimension (i.e., the same sequence length). Neural networks typically require inputs of a fixed size. By truncating excessively long signals and padding excessively short signals with zeros (or edge padding), data batches of uniform length can be generated, facilitating batch processing and training of the model.
[0040] In this embodiment, the length normalization submodule 213 uses a fixed time length or number of sampling points as a standard. For example, a standard length L = 1024 sampling points is set. For the amplitude-normalized signal of each channel, if its length is greater than L, the first L points are truncated from the signal start point; if its length is less than L, zeros are padded at the end of the signal until the length reaches L. After this step, the four-channel signal corresponding to each acoustic emission event is converted into a normalized tensor of size [4, L], which is the preprocessed time-domain signal with a uniform time length that is finally input to the feature extraction module.
[0041] Further, the feature extraction module includes an arrival time identification submodule and a time difference calculation submodule; the number of the plurality of acoustic emission sensors is four; the step of inputting the preprocessed time-domain signal into the feature extraction module to obtain the arrival time difference feature vector of the acoustic emission source of the structure to be monitored includes: Step b1: Input the preprocessed time-domain signal into the arrival time identification submodule to obtain the arrival times of the four channel signals.
[0042] Here, as Figure 3As shown, the feature extraction module 220 includes an arrival time identification submodule 221 and a time difference calculation submodule 222. Arrival time identification is the first step in feature extraction. Its core objective is to automatically and accurately detect the instant when the sound wave energy first arrives, i.e., the wave arrival time, from the preprocessed time-domain waveform of each sensor channel. This moment directly reflects the time it takes for the sound wave to travel from the sound source to the sensor, and it is the absolute benchmark for all subsequent geometric calculations and time difference analysis. Its accuracy fundamentally determines the final accuracy of the entire positioning system.
[0043] In this embodiment, the arrival time identification submodule 221 uses the Akaike Information Criterion (AIC) algorithm to automatically perform this critical task. This algorithm analyzes the signal of each channel, calculates the AIC function value of the signal through a sliding time window, and determines the point where the AIC function reaches its global minimum value as the initial time of sound wave arrival.
[0044] First, acquire the acoustic emission signals collected by four sensor arrays. The sensor coordinate matrix is defined as follows: ; Each acoustic emission event generates arrival time measurements for four channels, denoted as... ,in Represents the arrival of sound waves at the th The moment of each sensor.
[0045] Step b2: Input the arrival times of the four channel signals into the time difference calculation submodule to calculate the arrival time difference between every two acoustic emission sensors and combine them into a six-dimensional arrival time difference feature vector.
[0046] Here, the function of the time difference calculation submodule 222 is to derive the relative time difference, or time difference of arrival, between all possible sensor pairs based on four absolute arrival times. TDOA is the most critical observation in sound source localization because it directly eliminates the influence of the unknown absolute time of sound wave emission, retaining only the propagation time difference caused by differences in sound source location. For four sensors, six independent TDOA values can be calculated, which collectively encode the spatial location information of the sound source.
[0047] In this embodiment, the time difference calculation submodule 222 receives the arrival time array. The TDOA of all six sensor combinations (1-2, 1-3, 1-4, 2-3, 2-4, 3-4) is systematically calculated to form the original feature vector. To ensure that this feature has a stable numerical distribution when input into the subsequent neural network, the time difference calculation submodule 222 further standardizes it. This process uses a pre-stored mean vector calculated from the training dataset. and standard deviation vector According to the formula Perform calculations, where For a very small positive number (such as To prevent division by zero errors. The final output... This is the standardized six-dimensional time-of-arrival feature vector.
[0048] Furthermore, the fusion coding module is an encoder based on the Transformer architecture; the step of inputting the time difference of arrival feature vector into the fusion coding module to obtain the depth feature representation of the acoustic emission source includes: The arrival time difference feature vector is input into the encoder based on the Transformer architecture and encoded through the multi-head self-attention mechanism inside the encoder to obtain the deep feature representation; wherein, the encoder is used to extract the feature representation supporting the joint estimation of sound source and sound speed in the joint estimation model of the fusion Transformer and physical information neural network.
[0049] Here, the encoder based on the Transformer architecture is the core component for achieving deep feature fusion and abstraction in this method. Its key lies in the multi-head self-attention mechanism, which allows the model to dynamically calculate the correlation weights between any two elements (i.e., any pair of TDOA features) within the input time difference of arrival feature vector. This enables the encoder to go beyond simple sequential or local processing, globally capturing the time difference dependencies between all sensor pairs, thereby learning the complex spatial geometric patterns inherent in the six-dimensional TDOA features. Through this mechanism, the encoder transforms the raw, explicit physical observation features into a highly abstract, information-fusion-rich deep feature representation that intrinsically encodes all the high-order information needed for jointly inferring the sound source location and the sound velocity in the medium.
[0050] In this embodiment, the encoder employs a simplified Transformer encoder structure, including an embedding layer and position encoding: .in To embed the weight matrix, For positional encoding, specifically, the standardized six-dimensional TDOA feature vector is first projected into a higher-dimensional feature space (e.g., model dimension) through a linear embedding layer. =128). Subsequently, fixed-position encoding was added to incorporate sequence order information (here, feature order).
[0051] To better extract waveform features, a multi-head attention mechanism was used: .in, , , These represent the query, key, and value matrices, respectively. Specifically, the encoder core includes a multi-head self-attention layer where features are processed by calculating the query, key, and value matrices and then weighted by attention to achieve global interaction and information fusion. Furthermore, to capture the global dependencies between different sensors, this invention introduces a self-attention mechanism to overcome the limitations of local perception in traditional methods. The output of the attention layer is then non-linearly transformed through a feedforward neural network (e.g., two fully connected layers with a hidden layer dimension of 256, using the ReLU activation function) to enhance the model's expressive power. This design maintains sufficient representational power while controlling model complexity. Furthermore, to effectively mitigate the vanishing gradient problem in deep networks and accelerate model training, residual connections are used for each sub-layer, and normalization is applied. Ultimately, the output after processing by this encoder is a deep feature representation that can comprehensively characterize acoustic emission events and be used for subsequent multi-task prediction.
[0052] Furthermore, the multi-task output module includes a feature mapping submodule and a sound velocity analysis submodule; the step of inputting the depth feature representation to the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated medium equivalent sound velocity of the structure to be monitored includes: Step c1: Input the deep feature representation into the feature mapping submodule, and map the deep feature representation into three raw output values through the multilayer perceptron in the feature mapping submodule.
[0053] Here, as Figure 3 As shown, the multi-task output module 240 includes a feature mapping submodule 241 and a sound velocity resolution submodule 242. The feature mapping submodule 241 is the shared front end of the entire multi-task output module, and its core is a multilayer perceptron (MLP). The MLP's role is to decode and map the highly abstract deep feature representation generated by the fusion encoding module into three raw scalar values that are closer to the final output target. These three raw values are intermediate outputs for generating two-dimensional position coordinates and the medium-equivalent sound velocity; they have not yet undergone final physical constraints and range limitations.
[0054] In this embodiment, the MLP in the feature mapping submodule 241 consists of two fully connected layers, with a ReLU activation function used to introduce nonlinearity between them. The MLP receives the deep feature representation (e.g., a 128-dimensional vector) output by the previous module, and after two layers of nonlinear transformation, outputs three raw scalar values without any post-processing, denoted as the first raw output value o1, the second raw output value o2, and the third raw output value o3, respectively.
[0055] Step c2: The first and second original output values from the three original output values are directly output as the estimated two-dimensional position coordinates.
[0056] Here, in the joint estimation task, the x and y coordinates of the sound source location are the direct regression targets. The first original output value o1 and the second original output value o2, after preliminary mapping by the feature mapping submodule 241, directly correspond to the desired coordinate scale. Therefore, this method directly outputs these two values as the final two-dimensional position coordinate estimates without the need for additional complex transformations.
[0057] In this embodiment, the multi-task output module 241 directly assigns o1 to the estimated horizontal coordinate value. Assign the value of o2 to the estimated value of the ordinate. Output two-dimensional coordinates It directly corresponds to the spatial location within the monitoring area.
[0058] Step c3: Input the third original output value among the three original output values into the sound speed analysis submodule. Through the Sigmoid function unit and linear mapping unit in the sound speed analysis submodule, the third original output value is converted into the medium equivalent sound speed estimate value within the preset physical sound speed range.
[0059] Here, unlike location coordinates, the sound velocity of the medium is a parameter with a defined physical range (e.g., typically between 1500 m / s and 4500 m / s for concrete). The sound velocity analysis submodule is specifically designed to map the third raw output value o3 (whose range is theoretically unbounded) to a reasonable physical interval to ensure the physical meaning of the prediction results. This module first compresses o3 to the (0, 1) interval using a Sigmoid function unit, and then scales this value to between a preset minimum and maximum sound velocity using a linear mapping unit, thereby generating a physically reasonable sound velocity estimate.
[0060] In this embodiment of the application, the sound speed analysis submodule processes the data according to the following formula o3: .in, It is the Sigmoid function. and These are the minimum and maximum possible sound velocities preset based on the physical properties of the monitored material (such as concrete). For example, in this embodiment, the following settings are provided: =1500 m / s, =4500 m / s. With this design, regardless of the value of o3, the final output speed of sound estimate is... All values are strictly constrained within a reasonable range of [1500, 4500] m / s, ensuring the physical feasibility of the prediction.
[0061] Furthermore, before acquiring the raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored, the joint estimation model fusing the Transformer and the physical information neural network is trained according to the following steps: Step d1: Prepare the training dataset; the training dataset contains multiple training samples, each of which includes the known real location coordinates of the acoustic emission source and the original waveform signal collected by the corresponding sensor array.
[0062] Here, constructing a high-quality training dataset is fundamental to model training. Each training sample must contain a complete "input-output" correspondence: the input is the actual multi-channel raw waveform signal acquired by the sensor array; the output consists of two corresponding true values: the position coordinates of the artificially triggered or known acoustic emission source in the monitoring coordinate system, and the true sound velocity of the medium under the current experimental conditions (which can be calculated from the known location and arrival time or measured independently). The dataset needs to cover as many locations as possible within the monitoring area and the possible changes in sound velocity to ensure the completeness of model learning.
[0063] In this embodiment, four sensors at fixed positions are arranged on a 300mm×300mm concrete slab specimen, and multiple lead-breaking tests are conducted at different locations on its surface (e.g., a grid pattern) to simulate acoustic emission events. Each test records four channels of raw waveforms, and the coordinates of the lead-breaking point are precisely recorded as the true location label. Using the known sensor coordinates, the true sound source location, and the measured precise arrival time, the equivalent sound velocity of the medium (concrete slab) under that test can be calculated as the true sound velocity label. To ensure the quality and reliability of data acquisition and to provide accurate and effective acoustic emission event samples for subsequent model training, this embodiment systematically configures and optimizes the parameters of the data acquisition system. These settings aim to capture the most accurate and effective acoustic emission signals while suppressing interference from environmental noise and signal reflection. The specific parameter configuration comprehensively considers the acoustic characteristics of concrete materials, sensor performance, and industry standard practices, and has been verified and fine-tuned through multiple pre-experiments. The key system acquisition parameter settings are shown in Table 1 below.
[0064] Table 1 Specimen and Data Acquisition Parameter Settings
[0065] After acquiring the data, data augmentation operations were performed to ensure sufficient data for model training. These operations mainly involved normalization, data pruning, and scaling, expanding the data to 20 times the original size. The data was then divided into training and test sets in a 3:1 ratio. This resulted in the construction of a training dataset containing a large number of samples.
[0066] Step d2: Construct an initial joint estimation model; the initial joint estimation model includes the signal preprocessing module, the feature extraction module, the fusion coding module, and the multi-task output module.
[0067] Please see here. Figure 4 , Figure 4 This is a schematic diagram illustrating the training of the initial joint estimation model in an embodiment of this application. For example... Figure 4 As shown, the initial joint estimation model is based on Figure 2 The model is built using the architecture defined in [the document], forming a complete, end-to-end trainable neural network. Its input is the original waveform signal, and its output is the predicted sound source location and the speed of sound in the medium. The various modules in the model (signal preprocessing, feature extraction, fusion encoding, and multi-task output) are connected sequentially, and the parameters of each module are randomly initialized at the start of training.
[0068] In this embodiment of the application, based on the specific architecture of the method of this application, an initial joint estimation model fusing Transformer and PINN is built using deep learning frameworks such as PyTorch. This model fully integrates the signal preprocessing module (including sub-modules), feature extraction module (including sub-modules), Transformer-based fusion encoding module, and multi-task output module (including sub-modules) as described above.
[0069] Step d3: Construct a multi-objective loss function; the multi-objective loss function consists of a data loss term, a physical constraint loss term, and a boundary constraint loss term.
[0070] Here, the loss function is the core guiding the direction of model training. This application innovatively designs a composite loss function that simultaneously optimizes three objectives: 1) Data loss term: ensuring that the model's predicted values are as close as possible to the true labels in the training data, which is the foundation of supervised learning; 2) Physical constraint loss term: introducing the physical laws of sound wave propagation as soft constraints, forcing the model's prediction results (location and sound speed) to satisfy basic physical equations (such as the relationship between path difference and time difference), ensuring the physical rationality of the prediction even in cases of missing data or high noise; 3) Boundary constraint loss term: utilizing prior spatial information of the monitoring area (such as structural boundaries), constraining the predicted sound source location to fall within a reasonable physical space, avoiding physically impossible location predictions.
[0071] In the embodiments of this application, a multi-objective loss function is constructed. The specific form is as follows: .in, It is a loss based on the difference between predicted coordinates and true coordinates (such as L1 loss). It is based on the physical equation of sound wave propagation, comparing the theoretical TDOA calculated by predicting parameters with the actual measured TDOA loss; It penalizes the loss caused by predicted coordinates exceeding the monitoring area (0-300mm); and It is a hyperparameter used to balance the importance of various losses.
[0072] Step d4: Iteratively train the initial joint estimation model using the training dataset. Here, model training is an iterative optimization process. In each iteration, training data is input into the model, the loss is calculated, and the model parameters are updated through the backpropagation algorithm, gradually reducing the total loss and improving the model's predictive ability. This process continues until the model performance stabilizes on an independent validation set or meets a preset stopping criterion.
[0073] In this embodiment, iterative training is performed using mini-batch stochastic gradient descent. Each time, a small batch of samples is sampled from the training dataset, and the following steps of forward propagation, loss calculation, backpropagation, and parameter update are performed.
[0074] Each training iteration includes the following steps: Step e1: Input the original waveform signal of the current training sample into the initial joint estimation model, and process it sequentially through the signal preprocessing module, feature extraction module, fusion coding module and multi-task output module to obtain the predicted two-dimensional position coordinates and the predicted medium equivalent sound velocity.
[0075] This is the forward propagation phase of the training process. A small batch of raw waveform data is input into the model with the current parameter state. The data is processed by each module in the pipeline sequence defined by the model: first, it is preprocessed; then, TDOA features are extracted; next, the Transformer encoder performs feature fusion and abstraction; finally, the multi-task output head generates the predicted coordinates. and speed of sound .
[0076] In this embodiment of the application, for a batch of four-channel raw waveform data, the model automatically performs preprocessing (filtering, normalization, length standardization), feature extraction (AIC picking, TDOA calculation and standardization), fusion encoding (Transformer encoding), and multi-task output (MLP mapping and sound velocity analysis), and finally outputs the predicted position matrix and predicted sound velocity vector corresponding to the batch of samples.
[0077] Step e2: Calculate the data loss term based on the predicted two-dimensional position coordinates and the actual position coordinates of the current sample.
[0078] Here, the data loss term measures the model's performance in terms of pure data fitting ability. It directly calculates the error between the model's predicted sound source location and the actual location labeled in the dataset. To ensure model optimization, the loss should generally be as small as possible. The embodiments in this application use mean absolute error (L1 loss), and minimizing the data loss term forces the model to learn the direct mapping relationship from waveform features to location coordinates.
[0079] In this embodiment of the application, data loss item Calculated using the mean absolute error (MAE): ; in Given the batch size, sum the results across all samples within the batch. This directly reflects the average deviation of the predicted location.
[0080] Step e3: Based on the predicted two-dimensional position coordinates, the predicted medium equivalent sound velocity, and the known fixed sensor coordinates, and according to the physical equation of sound wave propagation, the predicted medium equivalent sound velocity is used as a variable to calculate the theoretical arrival time difference, and compared with the arrival time difference feature obtained by the feature extraction module for the current sample to calculate the physical constraint loss term.
[0081] This is the core concept of the Physical Information Neural Network (PINN). The loss term does not rely on the actual sound speed label, but instead uses physical principles to construct constraints. Specifically, it uses the sound source location currently predicted by the model. and speed of sound By combining the known sensor coordinates and based on the geometric relationship of sound wave propagation (distance difference = sound speed × time difference), the theoretical TDOA that should occur between each sensor pair is calculated. Then, this theoretical calculation is compared with the observed TDOA features actually calculated by the feature extraction module during the forward propagation of the current sample. Minimizing the difference between the two essentially forces the parameters (position and sound speed) learned by the model to satisfy basic physical laws, thereby embedding physical laws into the data-driven model and enhancing its interpretability and extrapolation capabilities.
[0082] In this embodiment of the application, the physical constraint loss term The calculation is as follows: .
[0083] The outer layer sums up the samples in the batch, while the inner layer sums up the samples from all sensor pairs. ; The observed TDOA is obtained from the feature extraction module; and It uses predicted coordinates The distances from the predicted sound source to sensors m and n are calculated using the known sensor coordinates.
[0084] Step e4: Calculate the boundary constraint loss term based on the predicted two-dimensional position coordinates and the preset monitoring area spatial boundary input.
[0085] Here, the boundary constraint loss term utilizes prior spatial knowledge of the monitoring area. In most practical applications, the acoustic emission source must be located inside the structure or within a reasonable area covered by the sensor array. This loss term penalizes cases where the predicted coordinates exceed the preset boundary range (e.g., outside the edge of a concrete slab), thereby guiding the model to output spatially reasonable results and avoiding absurd predictions (such as locating the sound source far away from the structure).
[0086] In this embodiment of the application, since the experimental data is collected in a two-dimensional plane from the cement block and the coordinate range of the collected data is between (300, 300), the output of the model is restricted in order to make the output value of the model conform to the actual situation.
[0087] Boundary constraint loss term Defined as: .
[0088] in The function is defined as: .
[0089] Step e5: The data loss item, physical constraint loss item, and boundary constraint loss item are weighted and summed according to preset weights to obtain the total loss value.
[0090] Here, the total loss value is the ultimate goal guiding the model parameter updates. Since the three losses have different dimensions and optimization objectives, simply adding them together might lead to training being dominated by one of them. Therefore, it is necessary to introduce balancing hyperparameters (weights) for the physical constraint loss term and the boundary constraint loss term to control their relative importance in the total loss. By adjusting these hyperparameters, we can balance the model's fit to the data labels, its adherence to physical laws, and its satisfaction of spatial priors.
[0091] Step e6: Update the parameters of the initial joint estimation model according to the total loss value until the preset training stopping condition is met, and obtain the trained joint estimation model of the fusion Transformer and physical information neural network.
[0092] This step is the parameter optimization stage in model training. Its core is to use the total loss value calculated in step e5 to calculate the gradient of this loss value with respect to all trainable parameters of the model via backpropagation. Subsequently, based on this gradient information, an optimization algorithm is used to iteratively update the model parameters, aiming to continuously reduce the total loss value, thereby gradually bringing the model's predicted output closer to the true label and conforming to physical and boundary constraints. This iterative update process continues until a preset training stopping condition is reached (e.g., reaching the maximum number of iterations, or the model performance stabilizing on a reserved validation dataset). At this point, the training process terminates, and the saved model is the trained joint estimation model that can be used for practical non-destructive acoustic emission monitoring.
[0093] In this embodiment, the model is trained iteratively using a training dataset. In each round (or each mini-batch) of training, steps e1 to e5 are executed to calculate the total loss value under the current model parameters, and a parameter update is performed accordingly. Through numerous such iterations, the model parameters are continuously adjusted and optimized. Finally, when the training process stops according to preset conditions, the resulting model parameter state defines a performance-optimized, fully trained joint estimation model fusing the Transformer and the physical information neural network. This model has learned the complex mapping relationship from the original waveform to position and sound speed, and inherently conforms to the physical laws of sound wave propagation and the spatial constraints of the monitoring area.
[0094] Further, updating the parameters of the initial joint estimation model based on the total loss value until a preset training stopping condition is met, to obtain the trained joint estimation model fusing the Transformer and the physical information neural network, includes: Step f1: Using the adaptive moment estimation optimization algorithm, calculate the gradient based on the total loss value and update the parameters of the initial joint estimation model.
[0095] Here, the adaptive moment estimation optimization algorithm is an advanced iterative optimization algorithm for training deep neural networks. It calculates estimates of the first moment (mean) and second moment (uncentered variance) of the loss function with respect to each parameter and dynamically adjusts the learning rate for each parameter using these estimates. This method can handle sparse gradients and adapt to different parameter update scales, thus achieving more stable and faster convergence, and is particularly suitable for optimizing models with complex modules and multi-task losses as described in this application.
[0096] In this embodiment, the Adam optimizer is used to perform parameter updates. The optimizer adjusts all weights and bias parameters of the model according to its update rules, based on the gradient of the total loss value calculated by backpropagation and combined with the first and second momentum estimates of each parameter maintained internally. Key hyperparameters are set as follows: learning rate of 0.0001, first-order moment decay factor (beta1) of 0.9, and second-order moment decay factor (beta2) of 0.999.
[0097] Step f2: Before updating the parameters of the initial joint estimation model, the calculated gradient is subjected to norm clipping to restrict the norm of the gradient vector to within a preset threshold.
[0098] In deep neural network training, especially in models with complex structures or loss functions involving multiple tasks, the gradient vector calculated through backpropagation may experience a "gradient explosion" phenomenon, where the norm of the gradient becomes abnormally large. This can lead to excessively large parameter update steps, destroying features already learned by the model and even causing the training process to diverge. Gradient norm pruning is a preventative technique that checks the norm (e.g., L2 norm) of the joint vector of gradients of all parameters in the entire model before applying the optimizer to update parameters. If the norm exceeds a preset threshold, the gradient vector is proportionally reduced to make its norm equal to the threshold, thus ensuring that the step size of each parameter update is controllable.
[0099] In this embodiment, gradient pruning is performed before the update step of the Adam optimizer is invoked. Specifically, the L2 norm of the gradients of all trainable parameters is calculated. The pruning threshold is set to 1.0. If the calculated gradient norm is greater than 1.0, all gradients are scaled proportionally (divided by the norm) so that the scaled gradient norm is equal to 1.0; if the gradient norm is less than or equal to 1.0, it remains unchanged. The formula is as follows: ; This operation effectively ensures the numerical stability of the training process.
[0100] Step f3: Periodically input the validation dataset into the joint estimation model of the current state to calculate the validation loss.
[0101] Here, the validation dataset is a set of independent samples that is not used for parameter updates during training but is used to evaluate the model's generalization performance. It is crucial to periodically evaluate the current model performance on the validation set (e.g., after each training epoch). By calculating the model's loss on the validation set (validation loss), we can monitor the model's performance on unseen data. This is the primary basis for determining whether the model is overfitting (over-memorizing training data details and losing generalization ability) and when to stop training.
[0102] In this embodiment, approximately 25% of the total dataset is divided into a validation dataset, and the remainder is a training dataset. After each complete training epoch, all samples from the validation dataset are input into the joint estimation model of the current training state, forward propagation is performed to obtain predicted values, and the total loss value is calculated (using the same formula as the training time, but without backpropagation and parameter updates). This value is the validation loss for the current epoch.
[0103] Step f4: Monitor the changing trend of the verification loss. When the verification loss does not decrease after reaching a preset number of rounds, trigger the early stop mechanism to forcibly terminate the training process.
[0104] Here, early stopping is a regularization technique used to prevent model overfitting. Its core idea is that in the early stages of training, the validation loss typically decreases with increasing training epochs, indicating improved generalization ability. However, when the model begins to overfit the training data, the validation loss stops decreasing and may even start to rise. Early stopping continuously monitors changes in the validation loss. If the validation loss does not decrease to a new minimum within several consecutive training epochs, it is determined that the model performance has stopped improving, and continuing training will only exacerbate overfitting; therefore, the training process is forcibly terminated. This helps to obtain a model with optimal generalization ability and saves computational resources.
[0105] In this embodiment, the patience value for the early stopping mechanism is set to 50 rounds, with a maximum of 300 training rounds. During training, a counter and a variable recording the best validation loss are maintained. After each round, the current validation loss is compared with the historical best validation loss. If the current validation loss is lower, the historical best value is updated to the current value, and the counter is reset to 0; otherwise, the counter is incremented by 1. When the counter value reaches the patience value of 50, meaning the validation loss has not reached a new low for 50 consecutive rounds, early stopping is triggered, and the training process is forcibly terminated.
[0106] Step f5: The joint estimation model corresponding to the model parameters obtained at the end of training is used as the trained joint estimation model of the fusion Transformer and physical information neural network.
[0107] Here, training may terminate due to reaching the preset maximum number of epochs or triggering an early stopping mechanism. At this point, the model parameters are in a final state. Typically, the model parameters at the training termination point are not used directly. Instead, the set of model parameters that performed best on the validation set throughout the entire training process is selected. This is because the model corresponding to this set of parameters has the minimum validation loss on independent data, meaning its generalization ability is optimal and it is most likely to perform well in real-world applications. The network defined by this selected set of optimal parameters is the final, trained joint estimation model that can be deployed and used for inference.
[0108] In this embodiment, during training, a snapshot of the current model's parameters is saved whenever the validation loss reaches a new low. When the training process terminates (due to reaching the maximum number of iterations or triggering early stopping), the snapshot of model parameters corresponding to the historical best validation loss is loaded. The network defined by these parameters is the final, trained joint estimation model fusing the Transformer and the physical information neural network, which will be used for all subsequent practical acoustic emission nondestructive monitoring tasks.
[0109] This application provides a method for non-destructive acoustic emission monitoring, comprising: acquiring raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored; inputting the raw waveform signals into a signal preprocessing module of a trained joint estimation model fusing Transformer and Physical Information Neural Network to obtain a preprocessed time-domain signal; the joint estimation model further includes a feature extraction module, a fusion encoding module, and a multi-task output module connected in sequence; inputting the preprocessed time-domain signal into the feature extraction module to obtain the arrival time difference feature vector of the acoustic emission source of the structure to be monitored; inputting the arrival time difference feature vector into the fusion encoding module to obtain the depth feature representation of the acoustic emission source; and inputting the depth feature representation into the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated equivalent sound velocity of the medium of the structure to be monitored. Thus, by fusing the physical constraints of Transformer and Physical Information Neural Network, the joint estimation of the sound source position and the sound velocity of the medium is achieved, improving the accuracy of non-destructive acoustic emission monitoring.
[0110] Based on the same application concept, this application also provides an acoustic emission non-destructive monitoring device corresponding to the acoustic emission non-destructive monitoring method provided in the above embodiments. Since the principle of the device in this application is similar to the acoustic emission non-destructive monitoring method in the above embodiments of this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0111] Please see Figure 5 , Figure 5 This is one of the functional block diagrams of an acoustic emission non-destructive monitoring device provided in an embodiment of this application. Figure 5 As shown, the acoustic emission non-destructive testing device 500 includes: The data acquisition module 510 is used to acquire the raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored.
[0112] The data processing module 520 is used to input the original waveform signal into the signal preprocessing module of the trained joint estimation model of the fusion Transformer and physical information neural network to obtain the preprocessed time domain signal; the joint estimation model also includes a sequentially connected feature extraction module, a fusion encoding module and a multi-task output module.
[0113] The feature extraction module 530 is used to input the preprocessed time-domain signal into the feature extraction module to obtain the arrival time difference feature vector of the acoustic emission source of the structure to be monitored.
[0114] The feature encoding module 540 is used to input the arrival time difference feature vector into the fusion encoding module to obtain the depth feature representation of the acoustic emission source.
[0115] The data generation module 550 is used to input the depth feature representation to the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated medium equivalent sound velocity of the structure to be monitored.
[0116] Further, the signal preprocessing module includes a bandpass filtering submodule, a signal normalization submodule, and a length normalization submodule; when the data processing module 520 inputs the original waveform signal into the trained joint estimation model of the fusion Transformer and physical information neural network to obtain the preprocessed time-domain signal, the data processing module 520 is specifically used for: The original waveform signal is input to the bandpass filter submodule to obtain the effective frequency band signal after filtering out environmental noise. The effective frequency band signal is input to the signal normalization submodule to obtain the amplitude-normalized signal. The amplitude-normalized signal is input into the length-normalization submodule to obtain the preprocessed time-domain signal with a uniform time length.
[0117] Furthermore, the feature extraction module includes an arrival time identification submodule and a time difference calculation submodule; the number of the plurality of acoustic emission sensors is four; when the feature extraction module 530 is used, the feature extraction module 530 is specifically used for: The preprocessed time-domain signal is input into the arrival time identification submodule to obtain the arrival times of the four channel signals; The arrival times of the four channel signals are input to the time difference calculation submodule to calculate the arrival time difference between every two acoustic emission sensors, and the results are combined into a six-dimensional arrival time difference feature vector.
[0118] Furthermore, the fusion encoding module is an encoder based on the Transformer architecture; when the feature encoding module 540 is used to input the arrival time difference feature vector into the fusion encoding module to obtain the depth feature representation of the acoustic emission source, the feature encoding module 540 is specifically used for: The arrival time difference feature vector is input into the encoder based on the Transformer architecture and encoded through the multi-head self-attention mechanism inside the encoder to obtain the deep feature representation; wherein, the encoder is used to extract the feature representation supporting the joint estimation of sound source and sound speed in the joint estimation model of the fusion Transformer and physical information neural network.
[0119] Furthermore, the multi-task output module includes a feature mapping submodule and a sound velocity analysis submodule; when the data generation module 550 is used to input the depth feature representation to the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated medium equivalent sound velocity of the structure to be monitored, the data generation module 550 is specifically used for: The deep feature representation is input into the feature mapping submodule, and the deep feature representation is mapped into three raw output values through the multilayer perceptron in the feature mapping submodule; The first and second original output values from the three original output values are directly output as the estimated two-dimensional position coordinates. The third original output value among the three original output values is input to the sound speed analysis submodule. The third original output value is converted into the medium equivalent sound speed estimate within the preset physical sound speed range by the Sigmoid function unit and the linear mapping unit in the sound speed analysis submodule.
[0120] Further, please refer to Figure 6, Figure 6 This is a second functional block diagram of an acoustic emission non-destructive monitoring device provided in an embodiment of this application. Figure 6 As shown, the acoustic emission non-destructive testing device 500 also includes: The data preparation module 560 is used to prepare the training dataset; the training dataset contains multiple training samples, each training sample including the known real location coordinates of the acoustic emission source and the corresponding raw waveform signal collected by the sensor array.
[0121] The model building module 570 is used to build an initial joint estimation model; the initial joint estimation model includes the signal preprocessing module, the feature extraction module, the fusion coding module and the multi-task output module.
[0122] The loss construction module 580 is used to construct a multi-objective loss function; the multi-objective loss function consists of a data loss term, a physical constraint loss term, and a boundary constraint loss term.
[0123] Model training module 590 is used to iteratively train the initial joint estimation model using the training dataset, wherein each training iteration includes the following steps: The original waveform signal of the current training sample is input into the initial joint estimation model, and processed sequentially through the signal preprocessing module, feature extraction module, fusion coding module and multi-task output module to obtain the predicted two-dimensional position coordinates and the predicted medium equivalent sound velocity. The data loss term is calculated based on the predicted two-dimensional position coordinates and the actual position coordinates of the current sample. Based on the predicted two-dimensional position coordinates, the predicted medium equivalent sound velocity, and the known fixed sensor coordinates, the predicted medium equivalent sound velocity is used as a variable to calculate the theoretical arrival time difference according to the physical equation of sound wave propagation. This theoretical arrival time difference is then compared with the arrival time difference feature obtained by the feature extraction module for the current sample to calculate the physical constraint loss term. The boundary constraint loss term is calculated based on the predicted two-dimensional position coordinates and the preset monitoring area spatial boundary input. The data loss term, physical constraint loss term, and boundary constraint loss term are weighted and summed according to preset weights to obtain the total loss value. Based on the total loss value, the parameters of the initial joint estimation model are updated until a preset training stopping condition is met, thereby obtaining the trained joint estimation model that fuses the Transformer and the physical information neural network.
[0124] Furthermore, when the model training module 590 updates the parameters of the initial joint estimation model based on the total loss value until a preset training stopping condition is met to obtain the trained joint estimation model fusing the Transformer and the physical information neural network, the model training module 590 is specifically used for: The adaptive moment estimation optimization algorithm is used to calculate the gradient based on the total loss value and update the parameters of the initial joint estimation model; Before updating the parameters of the initial joint estimation model, the calculated gradient is subjected to norm clipping to limit the norm of the gradient vector to within a preset threshold. The validation dataset is periodically input into the joint estimation model of the current state to calculate the validation loss; Monitor the changing trend of the validation loss. If the validation loss does not decrease after reaching a preset number of rounds, trigger the early stop mechanism to forcibly terminate the training process. The joint estimation model corresponding to the model parameters obtained at the end of training is used as the trained joint estimation model of the fusion Transformer and physical information neural network.
[0125] This application provides an acoustic emission nondestructive monitoring device, comprising: a data acquisition module for acquiring raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored; a data processing module for inputting the raw waveform signals into a signal preprocessing module of a trained joint estimation model fusing Transformer and Physical Information Neural Network to obtain a preprocessed time-domain signal; the joint estimation model further includes a sequentially connected feature extraction module, a fusion encoding module, and a multi-task output module; a feature extraction module for inputting the preprocessed time-domain signal into the feature extraction module to obtain an arrival time difference feature vector of the acoustic emission source of the structure to be monitored; a feature encoding module for inputting the arrival time difference feature vector into the fusion encoding module to obtain a depth feature representation of the acoustic emission source; and a data generation module for inputting the depth feature representation into the multi-task output module to simultaneously obtain an estimated two-dimensional position coordinate of the acoustic emission source and an estimated equivalent sound velocity of the medium of the structure to be monitored. Thus, by fusing the physical constraints of Transformer and Physical Information Neural Network, the joint estimation of the sound source position and the sound velocity of the medium is achieved, improving the accuracy of acoustic emission nondestructive monitoring.
[0126] Based on the same application concept, please refer to Figure 7 , Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 7 As shown, the electronic device 700 includes a processor 710, a memory 720, and a bus 730.
[0127] The memory 720 stores machine-readable instructions executable by the processor 710. When the electronic device 700 is running, the processor 710 and the memory 720 communicate through the bus 730. When the machine-readable instructions are executed by the processor 710, they perform the steps of the acoustic emission non-destructive monitoring method provided in the above embodiment. For specific implementation details, please refer to the method embodiment, which will not be repeated here.
[0128] Based on the same concept, this application also provides a computer-readable storage medium storing a computer program. When the computer program is run by a processor, it executes the steps of the acoustic emission non-destructive monitoring method provided in the above embodiments. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0129] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0130] In the embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some communication interfaces; indirect couplings or communication connections between devices or units may be electrical, mechanical, or other forms.
[0131] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0132] In addition, the functional units in the embodiments provided in this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0133] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0134] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. In addition, the terms "first", "second", "third", etc. are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0135] Finally, it should be noted that the above-described embodiments are merely specific implementations of this application, used to illustrate the technical solutions of this application, and not to limit them. The protection scope of this application is not limited thereto. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this application; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application. All should be covered within the protection scope of this application. Therefore, the protection scope of this application should be determined by the protection scope of the claims.
Claims
1. A method for non-destructive acoustic emission monitoring, characterized in that, The method includes: Acquire the raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored; The original waveform signal is input into the signal preprocessing module of the trained joint estimation model that integrates Transformer and physical information neural network to obtain the preprocessed time domain signal; the joint estimation model also includes a sequentially connected feature extraction module, a fusion encoding module and a multi-task output module; The preprocessed time-domain signal is input into the feature extraction module to obtain the arrival time difference feature vector of the acoustic emission source of the structure to be monitored; The arrival time difference feature vector is input into the fusion coding module to obtain the depth feature representation of the acoustic emission source; The depth feature representation is input to the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated medium equivalent sound velocity of the structure to be monitored.
2. The acoustic emission nondestructive monitoring method according to claim 1, characterized in that, The signal preprocessing module includes a bandpass filtering submodule, a signal normalization submodule, and a length normalization submodule; the signal preprocessing module, which inputs the original waveform signal into a trained joint estimation model fusing Transformer and physical information neural network, obtains a preprocessed time-domain signal, including: The original waveform signal is input to the bandpass filter submodule to obtain the effective frequency band signal after filtering out environmental noise. The effective frequency band signal is input to the signal normalization submodule to obtain the amplitude-normalized signal. The amplitude-normalized signal is input into the length-normalization submodule to obtain the preprocessed time-domain signal with a uniform time length.
3. The acoustic emission non-destructive monitoring method according to claim 1, characterized in that, The feature extraction module includes an arrival time identification submodule and a time difference calculation submodule; the number of the plurality of acoustic emission sensors is four; the step of inputting the preprocessed time-domain signal into the feature extraction module to obtain the arrival time difference feature vector of the acoustic emission source of the structure to be monitored includes: The preprocessed time-domain signal is input into the arrival time identification submodule to obtain the arrival times of the four channel signals; The arrival times of the four channel signals are input to the time difference calculation submodule to calculate the arrival time difference between every two acoustic emission sensors, and the results are combined into a six-dimensional arrival time difference feature vector.
4. The acoustic emission non-destructive monitoring method according to claim 1, characterized in that, The fusion coding module is an encoder based on the Transformer architecture; the step of inputting the time difference of arrival feature vector into the fusion coding module to obtain the depth feature representation of the acoustic emission source includes: The arrival time difference feature vector is input into the encoder based on the Transformer architecture and encoded through the multi-head self-attention mechanism inside the encoder to obtain the deep feature representation; wherein, the encoder is used to extract the feature representation supporting the joint estimation of sound source and sound speed in the joint estimation model of the fusion Transformer and physical information neural network.
5. The acoustic emission nondestructive monitoring method according to claim 1, characterized in that, The multi-task output module includes a feature mapping submodule and a sound velocity analysis submodule; the step of inputting the depth feature representation to the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated medium equivalent sound velocity of the structure to be monitored includes: The deep feature representation is input into the feature mapping submodule, and the deep feature representation is mapped into three raw output values through the multilayer perceptron in the feature mapping submodule; The first and second original output values from the three original output values are directly output as the estimated two-dimensional position coordinates. The third original output value among the three original output values is input to the sound speed analysis submodule. The third original output value is converted into the medium equivalent sound speed estimate within the preset physical sound speed range by the Sigmoid function unit and the linear mapping unit in the sound speed analysis submodule.
6. The acoustic emission nondestructive monitoring method according to claim 1, characterized in that, Before acquiring the raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored, the joint estimation model fusing the Transformer and the physical information neural network is trained according to the following steps: Prepare a training dataset; the training dataset contains multiple training samples, each training sample including the known real location coordinates of the acoustic emission source and the corresponding raw waveform signal collected by the sensor array; Construct an initial joint estimation model; the initial joint estimation model includes the signal preprocessing module, the feature extraction module, the fusion coding module, and the multi-task output module; Construct a multi-objective loss function; the multi-objective loss function consists of a data loss term, a physical constraint loss term, and a boundary constraint loss term. The initial joint estimation model is iteratively trained using the training dataset, wherein each training iteration includes the following steps: The original waveform signal of the current training sample is input into the initial joint estimation model, and processed sequentially through the signal preprocessing module, feature extraction module, fusion coding module and multi-task output module to obtain the predicted two-dimensional position coordinates and the predicted medium equivalent sound velocity. The data loss term is calculated based on the predicted two-dimensional position coordinates and the actual position coordinates of the current sample. Based on the predicted two-dimensional position coordinates, the predicted medium equivalent sound velocity, and the known fixed sensor coordinates, the predicted medium equivalent sound velocity is used as a variable to calculate the theoretical arrival time difference according to the physical equation of sound wave propagation. This theoretical arrival time difference is then compared with the arrival time difference feature obtained by the feature extraction module for the current sample to calculate the physical constraint loss term. The boundary constraint loss term is calculated based on the predicted two-dimensional position coordinates and the preset monitoring area spatial boundary input. The data loss term, physical constraint loss term, and boundary constraint loss term are weighted and summed according to preset weights to obtain the total loss value. Based on the total loss value, the parameters of the initial joint estimation model are updated until a preset training stopping condition is met, thereby obtaining the trained joint estimation model that fuses the Transformer and the physical information neural network.
7. The acoustic emission nondestructive monitoring method according to claim 6, characterized in that, The step of updating the parameters of the initial joint estimation model based on the total loss value until a preset training stopping condition is met, thereby obtaining the trained joint estimation model fusing the Transformer and the physical information neural network, includes: The adaptive moment estimation optimization algorithm is used to calculate the gradient based on the total loss value and update the parameters of the initial joint estimation model; Before updating the parameters of the initial joint estimation model, the calculated gradient is subjected to norm clipping to limit the norm of the gradient vector to within a preset threshold. The validation dataset is periodically input into the joint estimation model of the current state to calculate the validation loss; Monitor the changing trend of the validation loss. If the validation loss does not decrease after reaching a preset number of rounds, trigger the early stop mechanism to forcibly terminate the training process. The joint estimation model corresponding to the model parameters obtained at the end of training is used as the trained joint estimation model of the fusion Transformer and physical information neural network.
8. A non-destructive acoustic emission monitoring device, characterized in that, The acoustic emission non-destructive monitoring device includes: The data acquisition module is used to acquire the raw waveform signals collected by multiple acoustic emission sensors arranged on the structure to be monitored; The data processing module is used to input the original waveform signal into the signal preprocessing module of the trained joint estimation model that integrates Transformer and physical information neural network to obtain the preprocessed time-domain signal; the joint estimation model also includes a sequentially connected feature extraction module, a fusion encoding module and a multi-task output module; The feature extraction module is used to input the preprocessed time-domain signal into the feature extraction module to obtain the arrival time difference feature vector of the acoustic emission source of the structure to be monitored; The feature encoding module is used to input the arrival time difference feature vector into the fusion encoding module to obtain the depth feature representation of the acoustic emission source; The data generation module is used to input the depth feature representation into the multi-task output module to simultaneously obtain the estimated two-dimensional position coordinates of the acoustic emission source and the estimated medium equivalent sound velocity of the structure to be monitored.
9. An electronic device, characterized in that, include: The device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor communicates with the memory via the bus. The machine-readable instructions are executed by the processor to perform the steps of the acoustic emission non-destructive monitoring method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the acoustic emission non-destructive monitoring method as described in any one of claims 1 to 7.