Full-scale spatio-temporal fusion attention neural network with frequency domain enhancement and impact location method for composite sandwich panels

By using a frequency-domain enhanced full-scale spatiotemporal fusion attention neural network, the problems of insignificant signal features and low discriminability in impact positioning of composite sandwich panels were solved, and the accurate positioning of impacts on composite sandwich panels was achieved.

CN116050476BActive Publication Date: 2025-12-23CHONGQING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310054531.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-03
Publication Date
2025-12-23
Estimated Expiration
2043-02-03

Smart Images

  • Figure CN116050476B_ABST
    Figure CN116050476B_ABST
Patent Text Reader

Abstract

The application discloses a frequency domain enhanced full-scale space-time fusion attention neural network, comprising a backbone network and a frequency domain enhancement branch, wherein the backbone network comprises a full-scale hollow convolution module and a space-time fusion attention module; the output of the full-scale hollow convolution module is used as the input of the space-time fusion attention module; the features obtained by the backbone network and the frequency domain enhancement branch are fused to obtain the output of the neural network; the full-scale hollow convolution module is a unique three-layer convolution structure, the first convolution layer and the second convolution layer are one-dimensional hollow convolution layers, and the third convolution layer is a one-dimensional convolution layer; the space-time fusion attention module comprises a channel information aggregation unit, a channel interaction connection unit and a time correlation capturing unit; the frequency domain enhancement branch comprises an FEB module and an FFN module, and the signal is mapped to the frequency domain to consider the effect of the frequency component on impact positioning; and the application further provides a composite sandwich panel impact positioning method.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of deep learning, and particularly relates to a frequency domain enhanced full-scale spatio-temporal fusion attention neural network and a method for realizing composite sandwich panel impact positioning by adopting the neural network. BACKGROUND

[0002] With the increasing demand for lightweight and high-strength structures, carbon fiber composites are playing an increasingly important role in various equipment manufacturing fields. Composite sandwich panels (SCPs) are widely used in aerospace, shipbuilding, automotive and other fields due to their good designability, bending stiffness, compression stability and energy absorption characteristics, especially in panel structures such as aircraft wings, floors, ship hulls, etc. However, these structures are easily impacted by external foreign objects, such as runway stones, flying birds in the sky, floating ice in the water, space debris, etc., which may cause serious damage to the target structure, leading to a sharp decrease in the structural strength and load-carrying capacity of the SCP, posing a serious threat to the safe operation of the equipment. A large number of studies have shown that after the SCP is impacted, even if no surface damage is observed, its strength may decrease by 25%. Moreover, due to the unique structure of the composite sandwich panel, internal damage caused by impact is difficult to detect. Accurate positioning of the impact on the SCP is the premise of equipment service state monitoring and SCP damage detection.

[0003] Currently, common impact positioning technologies mainly include time difference positioning method, time reversal method, machine learning method, etc. The time difference positioning technology realizes impact positioning through the relationship between distance, angle and wave propagation time, and is a commonly used method in impact positioning. The time difference positioning technology has two key points: time difference and positioning algorithm. Now, time-of-flight (TOF) is often used as the time difference, but it has random errors caused by test equipment noise, temperature changes, etc., which are difficult to eliminate, resulting in that even if the positioning algorithm is perfect, accurate positioning cannot be completed. The time reversal method is a method of refocusing the sound source signal in space and time to realize positioning, wherein the response signal related to the sound source can be enhanced by controlling the reversal time and focusing conditions. However, this method usually needs to know the transfer function of the structure in advance through modeling method, which is very complex; and when the anisotropy of the composite material structure is obvious, the signal will not produce focusing effect, which seriously reduces the positioning effect of the time reversal method. The machine learning method first needs experts to design and select corresponding features from the signal according to prior knowledge, and then inputs these features into machine learning algorithms such as SVM and ANN to realize impact positioning. The effective selection of features is crucial to the positioning performance of the machine learning method, therefore, this method is highly dependent on expert experience and prior knowledge, and has unstable phenomenon in actual impact position identification.

[0004] With the development of industrial big data technology, a data-driven deep learning impact positioning method has emerged. This method directly inputs the original impact signal into the designed deep network, and directly outputs the positioning result after the network processes and analyzes the input signal. However, the current impact positioning method based on deep learning usually uses simple and shallow convolutional neural networks, recurrent neural networks and the like, which cannot well model the relationship between the impact signal and the impact position, and it is difficult to achieve satisfactory impact positioning accuracy. It can be seen that the current impact positioning method has many shortcomings, including error difficult to eliminate, positioning accuracy not enough, etc., which is difficult to be applied to engineering practice. SCP is generally composed of a panel, a bonding layer and a core material, as shown in Figure 1 The panel is generally a composite laminated plate with small density and thin thickness, and the core material is generally made of metal, non-metal foam or honeycomb. In addition, the bonding material (usually resin) combines the panel and the core material into a whole. Due to the special structure of the SCP, a part of the impact energy will be absorbed by the core material, so that the transmission of the impact signal in the sandwich panel is inhibited and the signal feature discrimination is reduced, and the impact positioning of the SCP is more difficult than that of ordinary metal and composite laminated plates. Therefore, it is difficult to obtain effective information by simply and independently analyzing each channel of the impact signal, and a more effective method must be explored to improve the discrimination of the signal features. SUMMARY

[0005] Therefore, the purpose of the present application is to provide a frequency domain enhanced full-scale spatiotemporal fusion attention neural network and a composite sandwich panel impact positioning method, wherein the neural network can extract all time scale features and spatiotemporal correlation of the impact signal, and improve the discrimination of the output features in a way that highlights effective information; the impact positioning method can position the impact on the composite sandwich panel.

[0006] To achieve the above purpose, the present application provides the following technical solutions:

[0007] The present application first proposes a frequency domain enhanced full-scale spatiotemporal fusion attention neural network, which includes a backbone network and a frequency domain enhancement branch. The backbone network includes a full-scale hollow convolution module and a spatiotemporal fusion attention module. The output of the full-scale hollow convolution module is used as the input of the spatiotemporal fusion attention module. The features obtained by the backbone network and the frequency domain enhancement branch are dynamically fused to obtain the output of the neural network.

[0008] The full-scale hollow convolution module is a three-layer convolution layer structure, the first convolution layer and the second convolution layer are one-dimensional hollow convolution layers, and the third convolution layer is a one-dimensional convolution layer; the convolution kernel composition of the one-dimensional hollow convolution layer follows the Goldbach conjecture rule, and a set of prime numbers is used as the size of the convolution kernel; by combining the convolution kernel size dimensions of the three one-dimensional convolution layers, the full-scale hollow convolution module can cover all the receptive fields of the input signal;

[0009] The space-time fusion attention module includes a channel information aggregation unit, a channel interaction connection unit and a time correlation capture unit; the channel information aggregation unit uses discrete cosine transformation to aggregate channel features, the channel interaction connection unit uses one-dimensional convolution to capture the hidden relationship between channels, and the time correlation capture unit fuses the internal attention mechanism and the gating mechanism to capture the time correlation of the input data;

[0010] The frequency domain enhancement branch includes an FEB module and an FFN module, which maps the signal to the frequency domain to consider the effect of frequency components on impact positioning.

[0011] Further, the convolution kernel size configuration of the full-scale hollow convolution module is as follows:

[0012]

[0013] Wherein, represents the convolution kernel size configuration of the i-th convolution layer, and: i∈{1,2} respectively corresponds to the first convolution layer and the second convolution layer, i=3 corresponds to the third convolution layer; {1,2,3,...,p k} represents all prime numbers from 1 to p k , and p k is determined depending on the length of the input signal.

[0014] Further, the receptive field of the full-scale hollow convolution module is represented as:

[0015]

[0016] Wherein, k i represents the convolution kernel size configuration of the i-th convolution layer; k 1 , k 2 and k 3 are the convolution kernel size configurations of the first convolution layer, the second convolution layer and the third convolution layer, respectively;

[0017] When the convolution kernel size of the third convolution layer satisfies the condition {k 3 =2,4}, there is:

[0018]

[0019] Wherein, denotes a set of positive even integers;

[0020] When the convolution kernel size of the third convolution layer satisfies the condition {k 3 = 1, 3}, there is:

[0021]

[0022] wherein, denotes a set of positive odd integers;

[0023] The receptive field of the full-scale dilated convolution module is represented as:

[0024]

[0025] wherein, represents a set of all positive integers within the length range of the input signal.

[0026] Further, in the process of performing the discrete cosine transform on the input by the channel information aggregation unit, the discrete cosine transform basis is generated according to the number of channels of the input component:

[0027]

[0028] wherein, denotes the generated discrete cosine transform basis; L denotes the length of the input component; m, n denote the data point sequence number in the input sequence;

[0029] The method for aggregating the input information by the channel information aggregation unit is:

[0030]

[0031] wherein, f denotes the scalar representing the channel obtained after the channel information aggregation; X m denotes the input sequence; freq denotes the discrete cosine transform basis selected from the discrete cosine transform basis set; LF denotes the low-frequency selection method.

[0032] Further, the channel interaction and the channel attention vector generation process of the channel interaction unit are as follows:

[0033] Iatt = sigmoid(conv1d(f))

[0034] wherein, Iatt denotes the spatial attention vector obtained after the channel information aggregation and interaction; conv1d denotes the one-dimensional convolution operation; sigmoid is the Sigmoid activation function; f denotes the scalar representing the channel obtained after the channel information aggregation.

[0035] Further, the principle of the time correlation capturing unit is as follows:​

[0036]

[0037] where Out denotes the output of the time correlation capturing unit; U denotes the vector generated by the linear layer from the original input signal; denotes the output feature containing time and spatial correlation; W o denotes the weight matrix; and:

[0038] U = Linear(X in )

[0039]

[0040] A = softmax(query x key), V = Linear(X in )

[0041] where X in denotes the original input signal; Linear denotes the linear layer; A denotes the attention map of the input information in time; softmax denotes the activation function; query and key denote the vectors obtained by performing twice vector scaling and offsetting operations on the spatial attention vector after linear transformation; V denotes the vector generated by the linear layer from the original input signal.

[0042] Further, the principle of the FEB module is:

[0043]

[0044] where FEB denotes the output of the FEB module; denotes the inverse Fourier transform; Padding denotes the zero padding operation; Q denotes the randomly reserved Fourier component; R denotes the matrix vector performing the frequency component selection; and:

[0045]

[0046] where Select denotes the random sampling method; denotes the Fourier transform; X in denotes the original input signal.

[0047] Further, the principle of the FFN module is:

[0048]

[0049] where FFN Swish denotes the output of the FFN module; Swish denotes the Swish activation function; Linear denotes the linear layer; X in denotes the original input signal; W denotes the weight matrix;

[0050] The Swish activation function is:

[0051] F Swish =x·sigmoid(x)

[0052] wherein F Swish represents the Swish activation function, sigmoid represents the sigmoid function, and x represents an input.

[0053] The application further provides an impact positioning method for a composite sandwich panel, comprising the following steps:

[0054] Step 1: data acquisition: impact any position on the composite sandwich panel to make the composite sandwich panel produce physical vibration after being impacted; the vibration information is captured by the sensors arranged on the composite sandwich panel and converted into an analog signal; an A / D converter is used to convert the analog signal into a digital signal; the above process is repeated to obtain an impact positioning data set;

[0055] Step 2: model construction: a frequency domain enhanced full-scale spatiotemporal fusion attention neural network as described above is constructed;

[0056] Step 3: model training: the neural network constructed is trained by using the impact positioning data set;

[0057] Step 4: impact positioning: the trained neural network is used for impact positioning.

[0058] Further, in the step 3, the Reduce LR On Plateau strategy is used to train the network, and when the error does not decrease for 50 epochs, the learning rate is adjusted to half of the learning rate of the previous step.

[0059] The application has the following beneficial effects:

[0060] The full-scale spatial-temporal fusion attention neural network with frequency domain enhancement of the application sets a full-scale cavity convolution module (OACM) and a spatial-temporal fusion attention module (STFAM) on the backbone network, wherein the OACM is used to extract information from the multi-channel impact signal and obtain full-scale features containing all the receptive fields, the OACM considers the long-term trend features while extracting the short-term features in the impact signal, so as to enhance the comprehensiveness of the features in representing the impact information; the STFAM includes a channel information aggregation unit, a channel interaction unit and a time correlation capturing unit, the time and spatial information of the multi-channel impact signal are considered together, the features rich in spatial-temporal correlation are obtained, and the important features in the signal are highlighted by assigning weights to the channels and signal elements, so that the discriminability of the impact signal features is improved; when applied to the composite sandwich panel, the problem of insignificant impact signal features and low discriminability caused by the energy absorption characteristics of the sandwich panel can be solved; in order to further excavate the hidden features of the signal that are difficult to be found in the time domain and effectively enhance the integrity of the output information, a frequency domain enhancement branch is also designed, the Encoder structure with a Swish gating linear network and a frequency enhancement block is adopted to process and analyze the impact signal from the frequency domain, the features that are not significant in the time domain but helpful for positioning can be extracted, and the positioning accuracy is improved; finally, the extracted feature information is dynamically fused to obtain the output of the neural network; when applied to impact positioning, the output of the neural network is output through an MLP layer to output the impact position; in summary, the full-scale spatial-temporal fusion attention neural network with frequency domain enhancement of the application improves the discriminability of the impact signal through a series of network module designs, so that the discriminability of the impact signal is effectively improved, and accurate impact positioning can be realized, such as accurate positioning of the impact on the composite sandwich panel. BRIEF DESCRIPTION OF DRAWINGS

[0061] In order to make the purpose, technical scheme and beneficial effects of the application more clear, the application provides the following drawings for illustration:

[0062] Figure 1 Fig. 1 is a structural schematic diagram of a composite sandwich panel; (a) carbon fiber reinforced polymer; (b) honeycomb sandwich composite panel; (c) corrugated sandwich composite panel;

[0063] Figure 2 Fig. 2 is a principle structure diagram of an embodiment of the full-scale spatial-temporal fusion attention neural network with frequency domain enhancement of the application;

[0064] Figure 3 Fig. 3 is a principle structure diagram of a full-scale cavity convolution module;

[0065] Figure 4 Fig. 4 is a principle structure diagram of a spatial-temporal fusion attention module;

[0066] Figure 5A principle structural diagram of a frequency domain enhancement branch;

[0067] Figure 6 A principle structural diagram of a data acquisition system;

[0068] Figure 7 A physical diagram of an impact positioning test platform;

[0069] Figure 8 Three evaluation index curves of a training process; (a) RMSE; (b) MAE; (c) R 2 _Score;

[0070] Figure 9 A positioning result diagram of an impact point;

[0071] Figure 10 A performance index column chart of the OSTFNet model of the present application and other four comparative models;

[0072] Figure 11 A positioning error curve diagram. DETAILED DESCRIPTION

[0073] The present application will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present application and implement it.

[0074] I. Frequency domain enhanced full-scale spatio-temporal fusion attention neural network (OSTFNet)

[0075] As shown in Figure 2 , the frequency domain enhanced full-scale spatio-temporal fusion attention neural network of the present embodiment includes a backbone network and a frequency domain enhancement branch. The backbone network includes an omni-scale atrous convolution module and a spatio-temporal fusion attention module. The output of the omni-scale atrous convolution module is used as the input of the spatio-temporal fusion attention module. The features obtained by the backbone network and the frequency domain enhancement branch are dynamically fused to obtain the output of the neural network.

[0076] 1.1. Omni-scale atrous convolution module (OACM)

[0077] The impact signal is a signal that decays rapidly over time, has different characteristics at different time scales, and these characteristics are very important for impact positioning tasks. In addition, effectively capturing the characteristics of different scales in the signal, such as short-term local characteristics and long-term trend characteristics, can enhance the feature's ability to represent the input information. To this end, the present embodiment is based on the Goldbach conjecture, that is, any even positive integer can be written as the sum of two prime numbers, and proposes an all-scale atrous convolution module (OACM) using a low-computational-resource-consuming atrous convolution network as the basic unit of the module to improve training efficiency. Based on this unique structural design, OACM can cover all receptive fields and extract all scale features in the impact signal, making full use of the effective information of the impact signal.

[0078] 1.1.1 Structure of all-scale atrous convolution module

[0079] As shown in Figure 3 , the all-scale atrous convolution module of the present embodiment is a three-layer convolutional layer structure, the first convolutional layer and the second convolutional layer are one-dimensional atrous convolutional layers, and the third convolutional layer is a one-dimensional convolutional layer. Specifically, the convolution kernel composition of the one-dimensional atrous convolutional layer follows the Goldbach conjecture rule and uses a set of prime numbers as the size of the convolution kernel; by combining the convolution kernel size dimensions of the three one-dimensional convolutional layers, the all-scale atrous convolution module OACM can cover all receptive fields of the input signal. In this way, for the impact signal with the characteristics of energy concentration and rapid amplitude decay, both short-term peak, energy concentration and other local significant information can be focused on, and long-term characteristics such as amplitude change speed can also be focused on at the same time, effectively improving the representation ability of the output feature and the network performance. Moreover, the present embodiment adopts dynamic design when determining the convolution kernel composition, that is, OACM can automatically determine the appropriate prime number combination according to the length of the input sequence, which is convenient to apply to tasks of different time sequences and has strong applicability.

[0080] In the present embodiment, the convolution kernel size configuration of the all-scale atrous convolution module is as follows:

[0081] Among them, represents the convolution kernel size configuration of the i-th convolutional layer, and: i∈{1,2} corresponds to the first convolutional layer and the second convolutional layer respectively, i=3 corresponds to the third convolutional layer; {1,2,3,...,p k} represents all prime numbers from 1 to p k , and p k is determined depending on the length of the input signal.

[0082] 1.1.2 Receptive field size of all-scale atrous convolution module

[0083] The receptive field is a basic concept of convolutional network, which represents that the unit in the feature depends on the size of the region in the input. The depth of the network has a huge impact on the receptive field of the convolutional layer. In the three-layer convolutional structure of OACM, each layer is a group of convolutions containing different kernel sizes. Therefore, the input signal will pass through multiple paths and different combinations of convolutional layers to finally obtain the output features containing all time scales.

[0084] The general formula for calculating the receptive field of one-dimensional convolution is as follows:

[0085]

[0086] wherein, represents the receptive field of the l+1th convolution layer, represents the receptive field of the lth convolution layer, k represents the kernel size of the l+1th layer, S l represents the step length of the lth convolution layer.

[0087] However, when processing impact signals, if the general convolution is used, the kernel will be very large due to the long sequence length, resulting in a decrease in training efficiency. Therefore, the method of the present embodiment uses a method of dilated convolution to improve the efficiency problem when processing long sequence data. The dilated convolution of the present embodiment actually increases the receptive field by expanding the kernel size. The size of the dilated convolution kernel corresponding to the general convolution is calculated as follows:

[0088] K c =S×(K o -1)+1

[0089] wherein, K c represents the size of the dilated convolution kernel, and K o represents the size of the general convolution kernel.

[0090] The receptive field of the OACM proposed in the present embodiment can be represented as follows when the step length is 1 and there is no pooling layer:

[0091]

[0092] wherein, k i represents the kernel size configuration of the i th convolution layer; k 1 , k 2 , and k 3 are the kernel size configurations of the first convolution layer, the second convolution layer, and the third convolution layer, respectively;

[0093] According to the Goldbach conjecture, {k 1 +k 2} can represent all positive even numbers, so {2(k 1 +k 2} represents positive even numbers starting from 4 with an interval of 4, {2(k 1 +k 2 )-2} represents positive even numbers starting from 2 with an interval of 4. Therefore, when the kernel size of the 3rd convolutional layer satisfies the condition {k 3 = 2, 4}, there is:

[0094]

[0095] wherein, represents a set of positive even numbers;

[0096] Similarly, {2(k 1 +k 2 )-1} represents positive odd numbers starting from 3 with an interval of 4, {2(k 1 +k 2 )-3} represents positive odd numbers starting from 1 with an interval of 4. Therefore, when the kernel size of the 3rd convolutional layer satisfies the condition {k 3 = 1, 3}, there is:

[0097]

[0098] wherein, represents a set of positive odd numbers;

[0099] In summary, the receptive field of the full-scale dilated convolution module is represented as:

[0100]

[0101] wherein, represents a set of all positive integers within the length range of the input signal. Therefore, by properly selecting the value of p k , the OACM proposed in the embodiment can cover all the receptive fields of the input and capture the full-scale features of the signal, enrich the features of the network and improve the impact positioning accuracy.

[0102] 1.2, Spatio-temporal Fusion Attention Module (STFAM)

[0103] As a kind of time series data, multi-channel impact signals also have rich time correlation, and have spatial correlation between the signal channels corresponding to multiple sensors. Capturing the complex time and spatial correlation of multi-channel signals is very important for impact positioning tasks, therefore, the embodiment proposes a spatio-temporal fusion attention module. As shown in FIG. 4, the STFAM is composed of two parts: a time attention module and a spatial attention module. Figure 4As shown, the space-time fusion attention module of the embodiment includes a channel information aggregation unit, a channel interaction connection unit, and a time correlation capturing unit. The channel information aggregation unit uses a discrete cosine transform to aggregate multi-channel information, the channel interaction connection unit uses a one-dimensional convolution to capture the hidden relationship between channels, and the time correlation capturing unit fuses an internal attention mechanism and a gating mechanism to capture the time correlation of the input data. In addition to being able to capture the space-time correlation of the impact signal, as an attention mechanism, the STFAM can also evaluate feature importance, highlight useful features, and suppress ineffective features to improve impact positioning accuracy. The input multi-channel signal is first subjected to a discrete cosine transform for channel information aggregation, and a one-dimensional convolution is used to capture the hidden relationship between channels; then the channel-weighted features are input into the internal attention block to capture the time correlation of the sequence, and the input is weighted after Softmax to obtain the attention output; at the same time, the STFAM also fuses a gating mechanism with three linear layers to control the output features; after these operations, the output with rich space-time correlation is finally obtained.

[0104] 1.2.1 Channel information aggregation unit

[0105] When aggregating information of input channels in the STFAM, instead of using a global average pooling (GAP) as in traditional channel attention models, the embodiment uses a discrete cosine transform that has good frequency energy aggregation capability. It can condense relatively important information in the input channel information together, and is very suitable for aggregating channel feature information.

[0106] Generally, the information within the input channel is diverse, and this is even more true for input with full-scale features extracted by the OACM. If GAP is simply used as the representation of channel information, the diversity of input features will be greatly suppressed, which is not conducive to model performance. The method of using a discrete cosine transform proposed in the embodiment can better express the original rich information of the features and preserve the diversity of the features, compared with the global average pooling.

[0107] In the process of discrete cosine transform of the channel information aggregation unit on the input, the discrete cosine transform bases are generated according to the number of channels of the input components:

[0108]

[0109] wherein, DCTbases generated; L represents the length of the input component; m, n represent the data point sequence number in the input sequence;

[0110] Then, according to the rule that the energy of the data after the discrete cosine transform is concentrated in the low-frequency part and the rule that CNN prefers low-frequency information, the low-frequency (LF) selection principle of the frequency components is determined, and the low-frequency frequency components are selected. Then, the freq is weighted with the input features by multiplication, and finally the aggregation of the input information is realized, and a group of scalar representing channel information containing rich input mode information is obtained. The method of the channel information aggregation unit for aggregating the input information is:

[0111]

[0112] wherein f represents the scalar representing the channel obtained after the aggregation of the channel information; X m represents the input sequence; freq represents the discrete cosine transform base selected from ; and LF represents the low-frequency selection method.

[0113] 1.2.2 Channel interaction unit

[0114] In the process of channel interaction, based on the group of scalars representing the channel obtained after the channel interaction, the STFAM learns the correlation between different channels by performing one-dimensional convolution operation thereon to interact the current channel with its adjacent channels, and obtains the importance of each feature channel.

[0115] Compared with the operation of realizing channel interaction by using the two-layer fully connected layer of “dimension reduction-dimension increase” in the traditional SE, the advantages of one-dimensional convolution are as follows: although the “dimension reduction-dimension increase” operation in SE reduces the model complexity and computational amount to a certain extent, it destroys the direct correspondence between the feature channel and the weight, which is not conducive to the generation of the channel attention map, while the one-dimensional convolution avoids this adverse effect; and the one-dimensional convolution realizes effective local cross-channel interaction instead of unnecessary global channel interaction, which is very reasonable from the perspective of the spatial arrangement of the sensors for collecting the impact signal, and also improves the computational efficiency.

[0116] Specifically, the channel interaction of the channel interaction unit and the process of generating the channel attention vector are as follows:

[0117] Iatt = sigmoid(conv1d(f))

[0118] wherein Iatt represents the spatial attention vector obtained after the aggregation and interaction of the channel information; conv1d represents the one-dimensional convolution operation; sigmoid is the Sigmoid activation function; and f represents the scalar representing the channel obtained after the aggregation of the channel information.

[0119] 1.2.3 Time correlation capturing unit

[0120] In the process of time correlation extraction, STFAM uses the operation of fusing the gating linear unit and the internal attention mechanism, and achieves the information interaction of the impact signal features by using the features captured by the spatial correlation to construct the query and key vectors, dynamically generates the weights of different connections, and then multiplies the vectors values obtained by the linear layer (in the gating linear unit) to complete the internal attention operation to capture the time correlation of the input data. The output features contain both time correlation and spatial correlation, and the output information is strictly controlled by combining the gating linear unit to obtain output features with rich spatio-temporal correlation.

[0121] First, the network performs a linear transformation on the spatial correlation containing features after channel information aggregation and interaction, and performs two simple per-dimension vector scaling and offset operations on the linearly transformed features to obtain query and key vectors. Second, the dot product operation is used as the scoring function to calculate the correlation of the impact signal features in the time scale, and then the softmax activation is performed to convert the output to the [0, 1] range to generate the attention map of the input information in time, and then the attention map is multiplied with the values vector generated by the linear layer to complete the scaling and weighting of the input data to obtain the output features containing time and spatial correlation. The operation process is as follows:

[0122] A = softmax(query x key), V = Linear(X in )

[0123]

[0124] where X in represents the original input signal; Linear represents a linear layer; A represents the attention map of the input information in time; softmax represents an activation function; query and key represent vectors obtained by performing two vector scaling and offset operations on the spatial attention vectors after linear transformation; V represents a vector generated by the original input signal through the linear layer.

[0125] Then, the gating linear unit is integrated into STFAM, where V has been integrated into STFAM as a vector in the internal attention, and U plays a role at the output position of the entire module to control the information output by STFAM. It is represented as follows:

[0126]

[0127] where Out represents the output of the time correlation capturing unit; U represents the values vector generated by the original input signal through the linear layer; represents the output feature containing time and space correlation; W o represents the weight matrix.

[0128] 1.3, frequency domain enhancement branch

[0129] In the backbone network, the two modules of OACM and STFAM extract the full-scale features of the impact signal and the time-space correlation in the time domain. However, as a fast-changing time series signal, the frequency domain features of the impact signal are still important for impact location prediction tasks. Therefore, this embodiment proposes to add a frequency domain enhancement branch outside the backbone network, as shown in Figure 5 The frequency domain enhancement branch includes an FEB module and an FFN module, which are used to map the signal to the frequency domain to consider the effect of frequency components on impact positioning. In addition, operating in the frequency domain can make the model better capture the global properties of time series.

[0130] FFEB adopts the Transformer structure with strong time series data processing capability, but this embodiment uses FEB to replace the multi-head attention module and analyzes and processes the impact signal in the frequency domain, and uses the powerful SwishGlu to replace the MLP network in the forward propagation module to improve the performance of the network, as shown in Figure 5 .

[0131] 1.3.1 FEB module (Frequency Enhanced Block)

[0132] As shown in Figure 5 , the FEB module first applies an MLP layer to linearly project the input impact signal, then performs discrete Fourier transform on the linearly projected signal to convert the signal from the time domain to the frequency domain, then randomly samples the obtained Fourier components to select a subset of components to appropriately represent the input information:

[0133]

[0134] where Select represents the random sampling method; represents the Fourier transform; X in represents the original input signal.

[0135] Then, multiply the selected Fourier components with the randomly initialized parameterized kernel-R; finally, pad the result of Q o R with 0 and perform IFT to convert the signal back to the time domain. The principle is:

[0136]

[0137] where FEB represents the output of the FEB module; represents the inverse Fourier transform; Padding represents the zero padding operation; Q represents the random reserved Fourier component; R represents the matrix vector performing the frequency component selection.

[0138] 1.3.2 FFN module

[0139] GLU is a powerful sequence modeling network, which can also be regarded as an MLP enhanced by gating mechanism, and it is very suitable for processing sequence data. As shown in Figure 5 , GLU is composed of two linear projections of component-wise multiplication, and one of the linear projections will use an activation function. In the classic GLU network, the activation function used is sigmoid, while the embodiment uses Swish activation function enhanced by self-gating mechanism. It has the characteristics of no upper bound, lower bound, smoothness, non-monotonicity, etc., which can avoid the situation that the gradient gradually approaches 0 and leads to saturation during slow training, and its smoothness can play a significant role in network optimization and generalization. Its function can be expressed as:

[0140] F Swish = x sigmoid(x)

[0141] where F Swish represents the Swish activation function; sigmoid represents the sigmoid function; x represents the input.

[0142] In SwishGLU, the input features are respectively obtained through two linear layers to obtain two intermediate vectors, and one of the vectors is activated using the Swish activation function; then, the two vectors are multiplied bit by bit, and finally, the output of the entire FFN Swish module is output through a linear layer. The principle of the FFN module is:

[0143]

[0144] where FFN Swis represents the output of the FFN module; Swish represents the Swish activation function; Linear represents the linear layer; X in represents the original input signal; W represents the weight matrix.

[0145] II. Impact positioning method of composite sandwich panel

[0146] The impact positioning method of the composite sandwich panel of the present application will be described below in combination with the specific implementation of the full-scale spatio-temporal fusion attention neural network enhanced in the frequency domain.

[0147] The impact positioning method of the composite sandwich panel of the present embodiment comprises the following steps:

[0148] Step one: data acquisition: impact any position on the composite sandwich panel to make the composite sandwich panel produce physical vibration after being impacted; the vibration information is captured by the sensors arranged on the composite sandwich panel and converted into analog signals; an A / D converter is used to convert the analog signals into digital signals; the above process is repeated to obtain an impact positioning data set, as shown in Figure 6 .

[0149] Step two: model construction: a full-scale spatio-temporal fusion attention neural network enhanced in the frequency domain is constructed. In order to achieve accurate positioning of the impact on the composite sandwich panel, the full-scale spatio-temporal fusion attention neural network enhanced in the frequency domain as described above is constructed in the present embodiment, and the multi-channel impact signals obtained from the impact signal acquisition system are directly input into the neural network. First, the OACM is used to extract information from the multi-channel impact signals to obtain full-scale features containing all the receptive fields, and the OACM considers both short-term features and long-term trend features in the extracted impact signals to enhance the comprehensiveness of the features in representing the impact information; then, the STFAM considers the time and space information of the multi-channel impact signals together by fusing channel attention, internal self-attention and a gating mechanism, obtains features rich in spatio-temporal correlation, and highlights important features in the signals by assigning weights to the channels and signal elements, so as to improve the discriminability of the impact signal features and solve the problem of low discriminability of the impact signal features due to the energy absorption characteristics of the sandwich panel; then, in order to further excavate hidden features of the signal that are difficult to be found in the time domain and effectively enhance the completeness of the output information, a frequency domain enhancement module is designed, which adopts an Encoder structure with a Swish gating linear network and a frequency enhancement block to process and analyze the impact signals from the frequency domain, so as to extract features that are not significant in the time domain but helpful for positioning and improve the positioning accuracy; finally, the extracted feature information is dynamically fused, and an MLP layer is used to output the impact position.

[0150] Step three: training the model: the neural network constructed is trained using the impact positioning data set. The ReduceLROnPlateau strategy is used to train the network, and when the error does not decrease for 50 epochs, the learning rate is adjusted to half of the learning rate of the previous step.

[0151] Step four: impact positioning: the trained neural network is used for impact positioning.

[0152] III. Experimental verification

[0153] The frequency domain enhanced full-scale space-time fusion attention neural network and the impact positioning method of the composite sandwich panel of the present embodiment are verified below in combination with specific examples.

[0154] 3.1, Signal acquisition and data set establishment

[0155] Specifically, the data acquisition system of the present embodiment is built with an impact positioning test bench, as shown in Figure 7 The impact positioning test bench includes an impact simulation device, a composite sandwich panel (SCP), a sensor, an A / D converter, a signal acquisition card, and data acquisition supporting software. First, the impact simulation device impacts any position on the composite sandwich panel, and the composite sandwich panel produces physical vibration under the impact. At the same time, the sensor arranged on the SCP captures the vibration information and converts it into an analog signal. Then, the analog signal is converted into a digital signal through sampling and quantization operations of the A / D converter. Finally, the digital signal is transmitted to the data acquisition software on the computer through the signal acquisition card for observation and storage of the impact signal waveform changes.

[0156] Table 1 Experimental equipment

[0157]

[0158] The detailed model of the experimental equipment used in the impact positioning test bench is shown in Table 1. The spring impact hammer is a semi-automatic impact device that can simulate impact behavior. The acceleration sensor can capture the physical vibration generated on the composite panel and convert the vibration into an analog electrical signal for transmission. The A / D converter can receive the analog signal from the sensor and convert it into a digital signal. The signal acquisition card is a connecting device between the data acquisition software on the computer and the signal, which can transmit the digital signal from the A / D converter to the data acquisition software to realize observation, storage, etc. of the impact signal.

[0159] In the test, the four ends of the composite sandwich panel are clamped by clamps on the test bench, as shown in Figure 7The size of the composite sandwich panel is 600x600mm, and the thickness is [0.5mm+3.0mm+0.5mm], with an aluminum honeycomb sandwich layer in the middle and a carbon fiber layer with a layering sequence of [0° / 90°] on the surface. For the convenience of the experiment, all impacts are applied to the surface within the designated area marked by a pen on the composite panel. The impact test area on the composite panel is a 450x450mm area in the middle, marked with lines at an interval of 30mm, resulting in a total of 212 nodes after removing the sensor-occupied area. Then, a spring impact hammer is used to perform multiple impacts on each impact node, and signal acquisition is performed at a sampling frequency of 1000Hz, resulting in a total of 8480 samples, each representing signal data with a sequence length of 1400 and a channel number of 8. After data preprocessing, the dataset is divided into a training set and a test set in a ratio of 7:3 and sent to the network for training. In addition, in addition to the aforementioned 212 impact points, 15 different impact points are randomly selected on the composite panel to evaluate the model performance and generalization.

[0160] 3.2 Model evaluation indicators and training

[0161] 3.2.1 Model performance evaluation indicators

[0162] This embodiment uses RMSE, MAE, and R 2 _Score to evaluate the model performance and positioning ability of the frequency domain enhanced full-scale spatio-temporal fusion attention neural network obtained by training. The smaller the values of RMSE and MAE, the stronger the positioning ability of the OSTFNet. In addition, the closer the value of R 2 _Score to 1, the higher the explanatory degree of the independent variable to the dependent variable, and the better the model performance. The calculation formulas of the three evaluation indicators are as follows:

[0163]

[0164]

[0165]

[0166] where Actual i represents the true value; Pred i represents the predicted value; and m represents the number of samples.

[0167] In addition, since the coordinates of each point are composed of (x, y) together, this embodiment also defines an error formula for impact positioning to quantify the distance between the predicted coordinates and the true coordinates, as shown below:

[0168]

[0169] where Error represents the impact positioning error; x actual and y actual represent the actual coordinates of the impact point; x pred and y pred represent the predicted coordinates of the impact positioning.

[0170] 3.2.2 Network training strategy

[0171] In terms of training optimization of the entire network, the Adam optimizer is used because it combines the advantages of Adagrad and RMSProp, uses the same learning rate for each parameter in the network, and can adapt independently as learning progresses. In addition, it also optimizes the model by utilizing historical information of gradients.

[0172] During the model training process, the learning rate is one of the most influential hyperparameters on performance, and appropriate learning rate selection has a positive effect on network training, model convergence, and maintaining model stability. In order to ensure both training speed and effectiveness, the initial learning rate is set to 0.001, and the Reduce LR On Plateau strategy is used to train the network. When the error does not decrease for 50 epochs, the learning rate is adjusted to half of the previous learning rate. In this way, the network can ensure training efficiency and model optimization: that is, in the initial stage of network training, the learning rate is larger to speed up the convergence of the network; when the model learns to a certain extent, the learning of the model to the data distribution tends to be stable, and a smaller learning rate is used to make the model approach the optimal point. In addition, the batch_size size in the training process is set to 64, and the dropout parameter in the network model is set to 0.2.

[0173] 3.2.3 Implement

[0174] The OSTFNet network model proposed in this embodiment is written using the pytorch deep learning framework based on Python 3.8 version; all experiments involved in this embodiment are performed on a Linux server equipped with Nvidia GTX GPU 3080.

[0175] 3.3 Model performance verification experiment

[0176] This embodiment uses the impact positioning dataset to test the impact positioning performance of OSTFNet, which shows excellent positioning effect of the model; in addition, the effectiveness of each module constituting OSTFNet is verified.

[0177] 3.3.1 Model performance verification

[0178] To verify the effectiveness of the proposed network, the model was trained and tested using the impact location dataset. Figure 8 (a)-(c) show the changes of RMSE, MAE, R 2 Score values during the training process of the model proposed in this embodiment. The network was trained for 1000 epochs. From the figure, it can be seen that in the early stage of training, the values of the three evaluation indicators change rapidly, which indicates that the full-scale features extracted by the network have strong discriminability and can well represent the impact information; in the later stage of training, the values of RMSE, MAE and R 2 Score tend to be stable, which keeps the positioning performance of the model stable, which can also be confirmed in the later positioning result figure. In addition, this "fast" and "stable" of model training also proves the correctness of the training strategy.

[0179] It can also be found that whether in the early rapid convergence stage or in the later stable stage of training, the coincidence degree of the training error curve and the test error curve is very high, which fully shows that the model proposed in this embodiment has good generalization. In the later stable stage, the average RMSE value of the model is 5.82, which indicates that in the prediction of impact coordinates, the average error between the predicted value and the true value is kept at a very low order of magnitude level, which is very excellent; and the average R 2 Score value reaches the level of 0.998, indicating that the regression fitting effect of the model is very good. All the above proves the superior positioning performance of the model proposed in this embodiment and the strong ability in feature extraction and spatio-temporal correlation mining of time series signals.

[0180] Finally, for the 15 impacts randomly performed on the composite plate, the positioning results of the OSTFNet network proposed in this embodiment are shown in Figure 9 . Its positioning accuracy is satisfactory, with an average error value of 5.157 mm, a minimum error value of 1.145 mm, and a maximum error value of 7.410 mm, which further proves that the OSTFNet of this embodiment has good ability in locating the impact position on the composite sandwich plate.

[0181] 3.3.2 Validation of the effectiveness of each module of the model composition

[0182] To evaluate the effectiveness of each module, part and improvement strategy in the OSTFNet proposed in this embodiment, after modifying the corresponding modules and improvement strategies, the network was trained and tested, and the experimental results are shown in Table 2. The results show that among the modules that make up the network, OACM has the greatest impact on the positioning performance of the network, followed by STFAM, and FEM has the least impact on the performance of the model; at the same time, various improvement strategies also have a certain impact on the positioning performance.

[0183] Firstly, it can be obviously seen that after replacing the OACM proposed in this embodiment with the common one-dimensional convolution, the RMSE and MAE errors in the network test process are significantly larger, the R 2 The Score value is also smaller, and the average positioning error of the network for the random 15 impact points reaches 16.983 mm, which is much higher than the OSTFNet proposed in this embodiment. This shows that, compared with the common convolution with fixed and local receptive field, the method of capturing all time scale features of the signal by the OACM module proposed in this embodiment can better represent the impact signal with energy concentration and rapid amplitude change, and has a very large advantage in modeling the characteristics of the impact signal.

[0184] In the STFAM, using GAP instead of the discrete cosine transform to aggregate the channel information respectively makes the RMSE, MAE errors rise by 2.055 and 1.643, and the average positioning error increases by 0.787 mm. This shows that, compared with GAP, using the discrete cosine transform can better aggregate the channel information and obtain a better channel feature representation. In addition, after removing the gating mechanism, the evaluation indicators of the model all have obvious changes, among which the positioning error rises by 1.767 mm, which shows that the gating mechanism widely used in the network processing long sequence data is very effective. After removing the STFAM module, the network performance evaluation indicators and the positioning error are greatly deteriorated, which shows that it is very necessary to use the STFAM to extract the space-time correlation of the multi-channel impact signal to improve the signal feature discrimination and positioning accuracy. In addition, using the full connection layer instead of the convolution in the channel interaction process has a relatively small impact on the model performance.

[0185] Finally, after removing the frequency domain enhancement module, the model positioning error significantly increases, the RMSE, MAE, R 2 The Score and other indicators are also greatly affected, which proves that processing the impact signal in the frequency domain can further mine the hidden features that are difficult to find in the time domain, effectively enhance the integrity of the output information, and improve the network positioning accuracy. In addition, after using the FFN instead of the SwishGLU, the model performance deteriorates and the positioning error increases, which fully proves the effectiveness of the SwishGLU.

[0186] Table 2 Model composition module performance experiment

[0187]

[0188] 3.4 Comparison test

[0189] This section compares the performance of the OSTFNet proposed in this embodiment with previous impact location networks and classic time series data processing networks. In addition, experiments are conducted to verify the applicability of the network to impacts of different energies.

[0190] 3.4.1 Comparison with classic networks

[0191] The OSTFNet network proposed in this embodiment is compared with the previous impact location networks SFNet and PZTNet. At the same time, based on the background that the impact signal is a kind of time series signal, the LSTM and RNN networks, which are recognized as having strong modeling ability in long sequence time series data processing, are also selected for comparison. Similarly, the same training strategy and hyperparameter selection are used for the five models.

[0192] In order to clearly analyze the differences between the five networks, Figure 10 the calculation and display of the three evaluation indicators RMSE, MAE, and R 2 _Score of the five network models after completing 1000 epoches training are shown in Table 1. From the table, it can be seen that the LSTM and RNN networks, which are good at sequence data processing, have better performance than SFNet and PZTNet in impact location, which shows that for the impact signal data set and the sequence length of 1400, the gating mechanism can well capture the long sequence features, which verifies the correctness of using the gating mechanism and considering the correlation of sequence of different scales in this embodiment. At the same time, the RMSE and MAE error values of the four comparison network models are significantly larger than those of the OSTFNet model, and the R 2 _Score value of the OSTFNet model is also higher than those of the other four networks, which shows that the features extracted by the OSTFNet network are more discriminative, making the network easily distinguish the differences between different impact locations. Moreover, the four comparison networks do not have unique network structure design for improving feature discriminability. This fully confirms that: the full-scale features including short-term detail features and long-term trend features extracted by the OACM module proposed and used in this embodiment can enhance the representation ability of features for impact information; the spatiotemporal correlation and inter-channel hidden connection captured by the STFAM module from multi-channel impact signals can effectively improve the discriminability of the signals; the way of analyzing and processing the impact signals in the frequency domain by the FEM module can extract time-domain insignificant features and other unique features, helping to improve the completeness of the output features for signal representation. It is because of the targeted design of these modules for improving the discriminability of impact signals that the OSTFNet model of this embodiment has better performance than the four comparison networks.

[0193] The positioning results of the network proposed in this embodiment and four comparative networks for 15 randomly selected impact points of composite plates are summarized in Table 3, wherein the maximum error, the minimum error, and the average error of positioning are recorded at the bottom of Table 3, and the best results are in bold. In addition, in order to better evaluate the positioning effect of the five networks, the positioning error values of the five networks for 15 random impacts are described in Table 4. Figure 10

[0194] As shown in Figure 11 , the OSTFNet proposed in this embodiment has only one error greater than SFNet and PZTNet in 15 impact positioning, which shows that the OSTFNet has better performance than the previous positioning network in the impact positioning task. Moreover, the error value of the OSTFNet for 11 impact positioning is less than that of the LSTM network, and the error of 14 impact positioning is less than that of the RNN network, which shows that the network proposed in this embodiment has a great advantage in long sequence data processing, and also proves the effectiveness of OACM in capturing all scale information. Overall, the OSTFNet has 10 positioning errors lower than all four comparative networks, and the 15 error changes are very stable and do not have the error fluctuations of other networks. This fully shows that under the joint action of OACM, STFAM, and FEM, the network output features with full-scale information, spatiotemporal correlation, and frequency domain information have good discriminability, which is specifically manifested in that the OSTFNet has excellent positioning performance and generalization ability.

[0195] According to the results in Table 3, it can be noted that compared with all the comparison methods, the OSTFNet has smaller maximum positioning error, minimum positioning error, and average positioning error. For 9 impacts in the X coordinate and 5 impacts in the Y coordinate, the OSTFNet of this embodiment has better positioning effect, followed by LSTM and RNN which are recognized as having advantages in time series data processing. The statistical results show that the OSTFNet has better positioning performance than other networks.

[0196] Table 3 Comparison of positioning performance of models

[0197]

[0198]

[0199] 3.4.2 Verification of applicability of OSTFNet under different energy levels

[0200] ​Considering the influence of the superior energy absorption characteristics of the composite sandwich panel on impact localization, three different impact energy levels (0.14J, 0.2J, 0.5J) were tested in this embodiment, and the collected impact signals were input into the OSTFNet to verify the localization effect and applicability of the OSTFNet to impact behavior of different energies, and the results are summarized in Table 3. As can be seen from Table 3, for 15 impact localizations with an energy of 0.5J, the OSTFNet still achieved very excellent results, with little difference from the aforementioned impact localization effect using 0.2J energy, and even achieved a lower minimum localization error. For 15 random impacts with an energy of 0.14J, although the maximum, minimum and average errors of the OSTFNet increased compared to the impact localization results with energies of 0.2J and 0.5J, it still achieved more accurate localization compared to the prior art. In general, although the vibrations generated by low-energy impacts are affected by the composite sandwich panel structure, causing some features in the impact signal to be suppressed and not highly distinguishable, with the help of the three targeted module designs, the output features of the OSTFNet are still highly distinguishable and achieve satisfactory localization effects.

[0201] Table 4 Comparison of localization effects of different energy impacts

[0202]

[0203] The above-described embodiments are merely preferred embodiments of the present application, and the scope of protection of the present application is not limited thereto. Any equivalent substitutions or transformations made by those skilled in the art based on the present application are within the scope of protection of the present application. The scope of protection of the present application is subject to the claims.

Claims

1. A frequency domain enhanced full-scale spatio-temporal fusion attention neural network, characterized in that: The frequency domain enhanced full-scale spatio-temporal fusion attention neural network comprises a backbone network and a frequency domain enhancement branch, the backbone network comprises a full-scale dilated convolution module and a spatio-temporal fusion attention module; The output of the full-scale dilated convolution module is taken as the input of the spatio-temporal fusion attention module; the features obtained by the backbone network and the frequency domain enhancement branch are dynamically fused to obtain the output of the neural network; The full-scale dilated convolution module is a three-layer convolution layer structure, the first convolution layer and the second convolution layer are one-dimensional dilated convolution layers, and the third convolution layer is a one-dimensional convolution layer; the convolution kernel of the one-dimensional dilated convolution layer is composed in accordance with the Goldbach conjecture rule, and a set of prime numbers is used as the size of the convolution kernel; by combining the convolution kernel size dimensions of the three one-dimensional convolution layers, the full-scale dilated convolution module can cover all the receptive fields of the input signal; The spatio-temporal fusion attention module comprises a channel information aggregation unit, a channel interaction connection unit and a time correlation capturing unit; The channel information aggregation unit adopts discrete cosine transformation to aggregate channel features, the channel interaction connection unit adopts one-dimensional convolution to capture the hidden relationship between channels, and the time correlation capturing unit fuses an internal attention mechanism and a gating mechanism to capture the time correlation of the input data; the input data is an impact positioning data set and is obtained by the following method: impacting any position on a composite sandwich panel to make the composite sandwich panel produce physical vibration after being impacted; the sensors arranged on the composite sandwich panel capture the vibration information and convert the vibration information into analog signals; an A / D converter is used to convert the analog signals into digital signals; the above process is repeated to obtain the impact positioning data set; The frequency domain enhancement branch comprises an FEB module and an FFN module, and the signal is mapped to the frequency domain to consider the effect of frequency components on impact positioning; The principle of the FEB module is as follows: wherein, represents the output of the FEB module; represents the inverse Fourier transform; represents the zero padding operation; represents the randomly reserved Fourier components; represents the matrix vector performing the frequency component selection; and: wherein denotes a random sampling method; denotes a Fourier transform; denotes the original input signal; The principle of the FFN module is as follows: wherein, represents an output of the FFN module; represents a Swish activation function; represents a linear layer; represents an original input signal; represents a weight matrix; The Swish activation function is as follows: wherein, represents a Swish activation function; represents a sigmoid function; represents an input.

2. The frequency-domain enhanced full-scale spatio-temporal fused attention neural network of claim 1, wherein: The convolution kernel size configuration of the full-scale dilated convolution module is as follows: wherein, denotes the number of convolutional kernels of the first convolutional layer, denotes the size configuration of the convolutional kernels of the first convolutional layer, and when i = 1 and j = 1, respectively, correspond to the first convolutional layer and the second convolutional layer, when i = 2 and j = 1, correspond to the third convolutional layer; denotes all prime numbers from 1 to and is determined depending on the length of the input signal.

3. The frequency-domain enhanced full-scale spatio-temporal fusion attention neural network of claim 2, wherein: The receptive field of the full-scale dilated convolution module is represented as follows: wherein, represents the first kernel size configuration of the layer convolutional layer; , and are kernel size configurations of the first, second and third convolutional layers, respectively. When the convolution kernel size of the 3rd convolution layer satisfies the condition , there is: wherein denotes the set of positive even integers; When the convolution kernel size of the 3rd convolution layer satisfies the condition has: wherein denotes the set of positive odd integers; The receptive field of the full-scale dilated convolution module is represented as follows: wherein represents the set of all positive integers in the range of the input signal length.

4. The frequency-domain enhanced full-scale spatio-temporal fusion attention neural network of claim 1, wherein: In the process of discrete cosine transformation of the input by the channel information aggregation unit, the discrete cosine transformation base is generated according to the number of channels of the input components: wherein denotes a generated discrete cosine transform basis; denotes a length of an input component; denotes a data point sequence number in an input sequence; the method for aggregating the input information by the channel information aggregation unit is: wherein, represents a scalar representing a channel after channel information aggregation; represents an input sequence; represents a discrete cosine transform basis selected from represents a low frequency selection method.​ 5. The frequency-domain enhanced full-scale spatio-temporal fusion attention neural network of claim 1, wherein: The channel interaction and channel attention vector generation process of the channel interaction connection unit are as follows: wherein, represents a spatial attention vector obtained after channel information aggregation and interaction; represents a one-dimensional convolution operation; is a Sigmoid activation function; represents a scalar representing a channel obtained after channel information aggregation.

6. The frequency-domain enhanced full-scale spatio-temporal fusion attention neural network of claim 1, wherein: The principle of the time correlation capturing unit is as follows: wherein, denotes the output of the time correlation capturing unit; denotes the vector generated from the original input signal by the linear layer; denotes the output feature containing time, spatial correlation; denotes the weight matrix; and: wherein, denotes the original input signal; denotes a linear layer; denotes the attention map of the input information over time; denotes an activation function; and denotes the vectors obtained by performing twice vector scaling and offset operations on the linearly transformed spatial attention vector; denotes the vector generated by the original input signal through the linear layer.

7. A method of impact location for a composite sandwich panel, characterized by: The method comprises the following steps: Step one: data acquisition: impacting any position on a composite sandwich panel to make the composite sandwich panel produce physical vibration after being impacted; the sensors arranged on the composite sandwich panel capture the vibration information and convert the vibration information into analog signals; an A / D converter is used to convert the analog signals into digital signals; the above process is repeated to obtain the impact positioning data set; Step two: model construction: constructing the frequency domain enhanced full-scale spatio-temporal fusion attention neural network according to any one of claims 1-6; Step three: model training: training the constructed neural network by using the impact positioning data set; Step four: impact positioning: using the trained neural network to perform impact positioning.

8. The composite sandwich panel impact location method of claim 7, wherein: In the third step, the network is trained using the Reduce LR On Plateau strategy. When the error does not decrease for 50 epochs, the learning rate is adjusted to half of the learning rate of the previous step.

Citation Information

Patent Citations

  • Impact positioning identification device and method

    CN111721450A

  • Electroencephalogram emotion recognition architecture based on time-space domain fusion and implementation method thereof

    CN113988129A