A multi-modal fusion method for low-speed impact positioning of composite laminated structures

Through the multimodal fusion low-speed impact positioning method of composite laminated structures, the moss growth optimization algorithm and ResNet18-NAM-Agent model are used to solve the problems of insufficient single-modal features and noise interference in composite laminated structures, and achieve high-precision and high-robustness impact positioning.

CN120277549BActive Publication Date: 2025-09-16BEIJING ZHONGTEST INFORMATION TECH DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510781458.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-16
Estimated Expiration
2045-06-12

AI Technical Summary

Technical Problem

Existing impact localization methods in composite laminated structures suffer from insufficient single-modal feature information, significant noise interference, and low efficiency of multimodal data fusion, resulting in insufficient positioning accuracy and robustness.

Method used

A low-speed impact localization method for composite laminated structures is adopted with multimodal fusion. The penalty coefficient of successive variational modal decomposition is dynamically optimized through the moss growth optimization algorithm. The envelope entropy criterion is combined to suppress noise and reconstruct the effective impact signal. The time-frequency spectrum is generated using continuous wavelet transform. A ResNet18-NAM-Agent multimodal regression model is constructed, and the channel-spatial attention mechanism is integrated to enhance the feature extraction capability. The dataset is expanded through a dynamic data augmentation strategy. The Huber loss function and the adaptive learning rate optimization model are combined to perform end-to-end impact position prediction.

Benefits of technology

It significantly improves the accuracy and robustness of impact positioning, solves the problems of insufficient feature extraction, spatial information utilization and noise resistance in traditional methods, and achieves high-precision spatiotemporal feature characterization of impact events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277549B_ABST
    Figure CN120277549B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal fusion low-speed impact positioning method for a composite laminate structure, which relates to the field of impact positioning technology. The method includes: obtaining impact response signals of multiple strain sensors on a composite laminate structure and combining them into a composite impact signal; using envelope entropy as a fitness function, and using a moss growth optimization algorithm to optimize the penalty coefficient of successive variational modal decomposition; performing successive variational modal decomposition on the composite impact signal based on the optimal penalty coefficient, and selecting some modal components with the largest energy distribution correlation coefficient to reconstruct a denoised signal; converting the denoised signal into a time-frequency spectrum, and converting the original composite signal into a two-dimensional coding diagram through a relative position matrix algorithm; constructing a multimodal sample data set including the time-frequency spectrum and the coding diagram; training a multimodal regression model and performing impact positioning. The present application scheme can significantly improve the modal purity of signal decomposition, enhance the characterization capability and fusion efficiency of multimodal features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of impact positioning, and in particular to a multi-modal fusion low-speed impact positioning method for a composite laminate structure. Background Art

[0002] Composite laminated structures are widely used in aerospace, wind power, and other fields due to their lightweight and high specific strength. However, they are extremely sensitive to low-velocity impacts, which can easily cause internal damage and threaten structural safety. Impact location technology, a core component of structural health monitoring, requires extracting effective features from sensor signals and mapping them to the impact location.

[0003] The core of the impact location method is to extract effective impact features from the signal data collected by the sensor and establish a mapping relationship between these features and the impact position. Commonly used impact location methods include the arrival time difference method, the reference database method, and data-driven machine learning methods. Among them, the arrival time difference method uses the geometric relationship between time difference, wave velocity and sensor position to achieve impact location, but for complex structures, it is difficult to accurately obtain the time difference and wave velocity in each direction; the reference database method matches a pre-established database containing impact features at different positions with the actual detection signal to achieve impact location. However, the establishment of the reference database requires a large amount of impact signal data, and the accuracy of positioning depends on the accuracy of the data; at present, the development of machine learning methods has improved the accuracy and robustness of impact location to a certain extent, but most machine learning methods only use single-dimensional information, are particularly susceptible to noise, and rely on feature engineering. In summary, the existing impact location methods have significant limitations in feature extraction, spatial information utilization, and noise resistance. Therefore, based on the above difficulties, the present invention proposes a multimodal fusion low-speed impact location method for composite laminated structures. Summary of the Invention

[0004] Purpose of the Invention

[0005] In order to solve the above problems, the purpose of the present invention is to provide a multimodal fusion low-speed impact positioning method for composite laminated structures, aiming to solve the problems of insufficient single-modal feature information, significant noise interference and low efficiency of multimodal data fusion in low-speed impact positioning of composite laminated structures, to achieve high-precision and high-robustness impact positioning, to overcome the traditional method's reliance on empirical parameter adjustment, and to enhance the ability to characterize the spatiotemporal characteristics of impact events in complex noise environments.

[0006] Technical Solution

[0007] To achieve the above objectives, the present invention provides a multimodal fusion low-speed impact positioning method for composite laminated structures. The method dynamically optimizes the penalty coefficient of successive variational modal decomposition based on the moss growth optimization algorithm, and combines the envelope entropy criterion to suppress noise and reconstruct the effective impact signal; generates a time-frequency spectrum through continuous wavelet transform to capture time-frequency features, and uses the relative position matrix algorithm to encode the signal sequence into a two-dimensional spatial distribution map; constructs a ResNet18-NAM-Agent multimodal regression model, integrates the channel-spatial attention mechanism to enhance feature extraction capabilities, and realizes cross-modal feature interaction through the agent attention mechanism; in addition, a dynamic data enhancement strategy is used to expand the data set, and the Huber loss function and adaptive learning rate are combined to optimize the model generalization performance, ultimately achieving end-to-end impact position prediction.

[0008] In a first aspect, the present invention provides a multi-modal fusion composite laminate structure low-speed impact positioning method, comprising:

[0009] Acquiring impact response signals of multiple strain sensors on the composite laminate structure and combining them into a composite impact signal;

[0010] The moss growth optimization algorithm is used to optimize the penalty coefficient of successive variational mode decomposition with envelope entropy as the fitness function.

[0011] The composite shock signal is subjected to successive variational modal decomposition based on the optimal penalty coefficient, and the modal components with the largest energy distribution correlation coefficient are selected to reconstruct the denoised signal.

[0012] The denoised signal is converted into a time-frequency spectrum through continuous wavelet transform, and the original composite signal is converted into a two-dimensional coding image through a relative position matrix algorithm;

[0013] Constructing a multimodal sample dataset comprising the spectrogram and coding graph;

[0014] A ResNet18-NAM-Agent multimodal regression model is trained and used to perform impact localization on the test set.

[0015] Furthermore, the parameters of the moss growth optimization algorithm include:

[0016] The population size is 10, the maximum number of iterations is 19, and the search range of the penalty coefficient is 500 to 60000.

[0017] Furthermore, the moss growth optimization algorithm includes a wind direction determination mechanism, a spore diffusion search, a double reproduction search, and a cryptobiotic mechanism, wherein the wind direction is dynamically adjusted by the positional relationship between the optimal individual and the majority of individuals in the population.

[0018] Furthermore, the update formula for the spore diffusion search is:

[0019]

[0020]

[0021]

[0022]

[0023]

[0024] Where, is the wind intensity attenuation factor; is the current evaluation number; is the maximum number of evaluations; for The number of mosses in the total number of moss individuals the proportion of Used to count the number of elements; It is a collection of most moss individuals; is the total number of moss individuals; is the distance spores travel under steady wind conditions; is a constant parameter 2; is the distance spores travel under turbulent wind conditions; For the moss individuals New mosses obtained by spore dissemination; It is a moss individual in the current moss population; For wind direction; is 0.2; 、 and A random number between 0 and 1.

[0025] Furthermore, the constraint criteria of the successive variational modal decomposition include minimizing the frequency domain bandwidth of the modal component, the spectrum overlap between the residual signal and the modal component, and the energy of the current mode near the historical modal center frequency.

[0026] Furthermore, the energy distribution correlation coefficient is calculated as follows:

[0027]

[0028] Where, is the energy distribution correlation coefficient; is the energy distribution of the modal component; is the energy distribution of the original signal; is the covariance; is the variance.

[0029] Furthermore, the relative position matrix algorithm includes:

[0030] After z-score normalization of the composite signal, the dimensionality is approximately reduced to m dimensions through segmented aggregation;

[0031] Construct an m×m matrix to represent the relative position relationship between timestamps and convert it into a grayscale coding image.

[0032] Furthermore, the construction of the multimodal sample dataset includes:

[0033] Gaussian, Poisson, salt and pepper, and multiplicative noise are added to the time-spectrogram, and the coded image is flipped, contrast enhanced, and histogram equalized, so that the data set is expanded to 5 times the original data volume.

[0034] Furthermore, in the multimodal regression model, ResNet18 is fused with the normalized attention mechanism, the proxy attention mechanism is used for multimodal feature fusion, the channel attention calculates the channel importance through the batch normalization weight, and the spatial attention generates the spatial weight through the channel mean;

[0035] The query and value of the proxy attention mechanism come from the time-frequency graph feature sequence, and the key comes from the encoding graph feature sequence. The fusion formula is:

[0036]

[0037]

[0038] Where, is the agent feature matrix; is global average pooling; , , , 、 and is the weight parameter, and is the original feature; is the feature sequence obtained after the fusion of the agent attention mechanism; is the softmax function; is the scaling factor; and is the position offset; It is a depth-wise separable convolution operation.

[0039] Furthermore, in the ResNet18-NAM feature extraction network, the NAM module of the time-spectrogram branch is inserted into the first two residual layers, and the NAM module of the coding graph branch is inserted into the first three residual layers and the post-fourth residual layer.

[0040] Furthermore, a dynamic dual-weight envelope entropy step is included to construct a composite fitness function by quantifying the energy concentration of the signal in the time domain and the sparsity in the frequency domain. The dual-weight envelope entropy formula is defined as:

[0041] Where, For the Dynamic double-weighted envelope entropy of modal components; is the total number of signal sample points; is the frequency domain sparse factor; is the time domain attenuation factor; is the time point of the signal; is the frequency point of the signal; For the The envelope normalized value of the modal component.

[0042] Furthermore, it also includes a spatiotemporal collaborative attention mechanism, which coordinates the cross-modal interaction strength by adjusting the temporal convolutional gating unit and the spatial dynamic weight. The spatiotemporal collaborative attention formula is defined as:

[0043]

[0044] Where, It is the spatiotemporal coordinated attention output; is the feature vector dimension; is channel-by-channel multiplication; It is the time domain convolution gating; is the spatial dynamic weight.

[0045] In a second aspect, the present invention further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, the aforementioned multi-modal fusion composite laminate structure low-speed impact positioning method is executed.

[0046] This method dynamically adjusts the penalty coefficients of successive variational modal decomposition based on a moss growth optimization algorithm, combines it with a dynamic dual-weight envelope entropy criterion to screen high-fidelity modal components, suppress noise, and preserve the transient characteristics of the impact. It captures time-frequency features by generating a time-frequency spectrum through continuous wavelet transform, and uses a relative position matrix algorithm to map the signal sequence into a two-dimensional coding diagram to explicitly represent the spatial distribution. It constructs a ResNet18-NAM-Agent multimodal regression model, introduces a cross-modal spatiotemporal collaborative attention mechanism to achieve deep interaction between time-frequency and spatial features, and combines dynamic data enhancement with an adaptive training strategy to optimize the model's generalization capabilities. This solution significantly improves the modal purity and stability of signal decomposition, enhances the representation capability and fusion efficiency of multimodal features, and thus achieves high-precision and robust impact localization in complex noisy environments. It systematically addresses the core issues of traditional methods, such as single features, noise sensitivity, and computational redundancy.

[0047] Beneficial effects

[0048] By implementing the multi-modal fusion composite laminate structure low-speed impact positioning method provided by the present invention, the following technical effects are achieved:

[0049] (1) This application optimizes the balance between noise suppression and impulse feature retention during signal decomposition by introducing a time-domain attenuation factor and a frequency-domain sparse factor to reconstruct envelope entropy calculation. This method significantly improves the purity and decomposition stability of the modal components and enhances the ability of the denoised signal to retain the transient characteristics of the impulse, thereby providing input data with a higher signal-to-noise ratio for the subsequent positioning model.

[0050] (2) By coordinating the interaction strength of multimodal features through temporal convolution gating and spatial dynamic weighting, the traditional attention mechanism solves the problem of insufficient modeling of correlation in the spatiotemporal dimension. This mechanism achieves efficient complementary fusion of time-frequency features and spatial distribution features, improves the model's robustness to complex noise and boundary reflection interference, and reduces the computational redundancy of cross-modal feature fusion.

[0051] (3) Through noise injection and image transformation operations, data diversity is enhanced for the physical characteristics of the time-spectrogram and the encoding graph, respectively, to construct a multimodal dataset with strong generalization capabilities. This strategy effectively alleviates the model's dependence on limited annotated data, suppresses overfitting during training, and ensures positioning consistency in sparse sensing or non-uniform noise environments.

[0052] (4) Through segmented aggregation approximation and relative position matrix construction, the one-dimensional signal sequence is mapped into a two-dimensional spatial distribution code map, explicitly characterizing the spatial propagation characteristics of the impact signal. This method makes up for the shortcomings of traditional time-frequency analysis methods in capturing spatial dimension information, providing a more comprehensive spatiotemporal feature input for the deep learning model, thereby improving the spatial resolution and positioning accuracy of the impact position. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to make the above-mentioned multi-modal fusion composite laminate structure low-speed impact positioning method of the present invention more obvious and easy to understand, the following will briefly introduce the drawings required for use in the specific implementation of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0054] Figure 1 It represents a flow chart of the present application method;

[0055] Figure 2 Schematic diagram of strain gauge pasting for the multimodal impact localization method based on ResNet18-NAM-Agent;

[0056] Figure 3 Schematic diagram showing a composite shock signal;

[0057] Figure 4 Schematic diagram showing the results of successive variational mode decomposition based on the moss growth algorithm optimization;

[0058] Figure 5 Schematic diagram showing the spectrum corresponding to the successive variational mode decomposition structure optimized based on the moss growth algorithm;

[0059] Figure 6 Schematic diagram showing the comparison between the reconstructed denoised shock signal and the original composite shock signal;

[0060] Figure 7 Schematic diagram of the multimodal impact localization model based on ResNet18-NAM-Agent. DETAILED DESCRIPTION

[0061] Example 1:

[0062] A multi-modal fusion composite laminate structure low-speed impact positioning method is provided, and the method process is as follows Figure 1 Shown, including:

[0063] Step 1: Acquire impact response signals of multiple strain sensors on a composite laminate structure specimen, and arrange the signals of the multiple sensors in the order of sensor 1, sensor 2, ..., sensor n to combine them into a composite impact signal;

[0064] like Figure 2 As shown in the figure, a 400mm×400mm square monitoring area was selected on a 650mm×650mm×3mm composite laminate specimen. Four strain sensors were symmetrically attached at the four corners of the monitoring area, with the strain sensor axes parallel to the diagonals of the square at the corners and at a 45° angle to the side edges of the monitoring area. A two-dimensional rectangular coordinate system was established for the monitoring area of ​​the composite laminate structure, with the lower left vertex of the monitoring area defined as the coordinate origin (0mm,0mm). The x-axis was defined as a horizontal line from the origin, and the y-axis as a line perpendicular to the x-axis from the origin. The monitoring area was divided into a grid of moderate size and uniform distribution, resulting in a total of 81 grid points. These 81 grid points were struck three times with an impact hammer, and 15 randomly selected points were struck once. The signal acquisition device was set to a sampling frequency of 10kHz to collect the impact signals collected by the strain sensors.

[0065] like Figure 3 As shown, the strain signals collected by the preferred four strain sensors are arranged on the time axis according to the arrangement order of the strain sensors to form a composite impact signal.

[0066] Step 2: Using envelope entropy as the fitness function, the moss growth optimization algorithm is used to find the penalty coefficient after the composite impact signal in step 1 is decomposed by the successive variational mode decomposition algorithm. ;

[0067] In finding the penalty coefficient for the optimal decomposition When the envelope entropy is selected as the fitness function of the moss optimization algorithm, the envelope entropy is an indicator of signal complexity, reflecting the uncertainty of signal energy in the envelope distribution. The energy of the impulse signal is concentrated in a few moments, the envelope distribution is more sparse and orderly, and the envelope entropy value is small. The energy distribution of the noise signal is uniform, the envelope distribution is disordered, and the envelope entropy value is large. Therefore, in the optimization process, we try to find the envelope entropy value that is small. The envelope entropy value is calculated as follows:

[0068]

[0069]

[0070] Where, for The normalized form of is the envelope of the intrinsic modal component after Hilbert transform; is the total number of signal sample points; is the envelope entropy value.

[0071] When initializing the relevant parameters of the moss growth optimization algorithm, the population size is set to 10, the maximum number of iterations is set to 19, and the penalty parameter is set to 10. The selection range of is 500 to 60,000. Each moss individual in the moss population represents a penalty parameter of the successive variational mode decomposition algorithm. The moss growth optimization algorithm includes four key stages: wind direction determination, spore diffusion search, double diffusion search and cryptobiotic mechanism.

[0072] The wind direction mechanism of moss propagation uses the positional relationship between the majority of moss individuals and the optimal moss individual to determine the evolutionary direction of all moss individuals in the population. This evolutionary direction can effectively help the moss growth optimization algorithm avoid falling into a local optimal solution. Its expression is:

[0073]

[0074]

[0075] Where, For each moss individual in the current population relative to The distance set of ; is the optimal solution in the current moss population; It is a moss individual in the current moss population; It is a collection of most moss individuals; For wind direction; for The total number of individuals in For the The distance between a moss individual and the optimal moss individual.

[0076] The spore diffusion search mechanism simulates the propagation characteristics of spores under both steady and turbulent wind conditions, allowing individuals to make random selections. This can prevent the fixed step size from converging slowly in the early stages and causing non-convergence in the later stages. Based on the propagation principle of moss spores, the population position is updated as follows:

[0077] The update formula for the spore diffusion search is:

[0078]

[0079]

[0080]

[0081]

[0082]

[0083] Where, is the wind intensity attenuation factor; is the current evaluation number; is the maximum number of evaluations; for The number of mosses in the total number of moss individuals the proportion of Used to count the number of elements; It is a collection of most moss individuals; is the total number of moss individuals; is the distance spores travel under steady wind conditions; is a constant parameter 2; is the distance spores travel under turbulent wind conditions; For the moss individuals New mosses obtained by spore dissemination; It is a moss individual in the current moss population; For wind direction; is 0.2; 、 and A random number between 0 and 1.

[0084] When random number When it is less than 0.8, the dual reproduction search mechanism is triggered. Unlike the traditional metaheuristic algorithm, this method increases the proportion of methods that only change a single dimension, enhances the overall local search capability, and further updates the moss population position with a certain probability according to the dual reproduction mechanism of moss. The expression is:

[0085]

[0086]

[0087]

[0088]

[0089] Where, Used to evaluate the optimal moss individual Whether the particles in the are utilized; For A random vector between 0 and 1 with the same dimension; 、 is a random number between 0 and 1, is 0.5, if , then the double reproductive search is simulated in the sexual reproduction stage, otherwise the double reproductive search is simulated in the vegetative reproduction stage; Used to control the wind direction of the individual Distance moved in a certain direction; For the moss individuals New mosses obtained by spore dissemination; for The particles, is a random number that does not exceed the maximum dimension of the individual; for The particles; for The particles.

[0090] The cryptobiotic mechanism optimizes the search process according to the cryptobiotic phenomenon of mosses, and retains the records of moss individuals generated during each iteration. Once the maximum number of records reaches 10 or the population iteration is completed, the cryptobiotic mechanism is triggered to restore the optimal moss individual and replace the current moss individual, thereby updating the population position and the optimal fitness function value, ensuring the ability of the entire population to conduct global search.

[0091] Step 3: Set the optimal penalty factor Apply successive variational modal decomposition to decompose the composite impulse signal in step 1 to obtain multiple modal components. Calculate the correlation coefficient between the energy distribution of the modal components and the energy distribution of the composite impulse signal in step 1. Select the first M modal components with the largest correlation coefficients for signal reconstruction to obtain the denoised impulse signal.

[0092] Successive variational mode decomposition (SVMD) performs continuous variational mode decomposition on the signal until the reconstruction error is less than a set threshold. The biggest advantage of SVMD is that it does not require the number of modes in the signal to be known in advance, which reduces computational complexity. SVMD has the advantages of strong adaptability, high decomposition accuracy, and good robustness, making it particularly suitable for feature extraction in impact signals and complex noise environments. SVMD is used to decompose the target composite impact signal data. The decomposition process specifically includes:

[0093] The target composite shock signal is input into the successive variational mode decomposition algorithm, and the original composite shock signal Decomposed into two signals, including the L-order natural mode component signal and the residual signal , where the residual signal Except The signal other than the one above contains the sum of all previously obtained modes and the unprocessed part of the composite impulse signal. , the expression is:

[0094]

[0095]

[0096] Where, For the L-order natural mode component signals.

[0097] To ensure that each modal component can be compact around its center frequency to ensure that its energy is concentrated in the frequency domain, the following criteria should be minimized for the Lth-order modal component:

[0098]

[0099] Where, It is an objective criterion for minimizing the bandwidth of the modal component in the frequency domain; For time The derivative of is the unit pulse function; is the imaginary unit in complex numbers; is a complex exponential function used to convert the signal into the frequency domain; is the center frequency of the L-th order natural mode component;

[0100] exist At the frequency containing the effective component, the residual signal should be The energy of is the smallest, so to minimize the residual signal and the Lth-order modal component The spectrum overlap is expressed as:

[0101]

[0102] Where, The constraint criterion is to minimize the spectral overlap between the residual signal and the Lth intrinsic mode component; For filter The impulse response of

[0103] By minimizing and Two constraint criteria are used to obtain the Lth modal component of the composite impulse signal. , but the Lth modal component obtained There may be significant modal aliasing between the previously obtained L-1 modal components. To avoid this, it is necessary to ensure that There should be less energy at frequencies near the center frequency of the modal component obtained previously, expressed as:

[0104]

[0105] Where, To ensure that the current modal component A constraint criterion of having minimum energy around the center frequency of the previously obtained natural mode components; The center frequency is filter;

[0106] The last constraint criterion is to ensure that all modal components and unprocessed parts can be completely reconstructed into the original composite impact signal data, which is expressed as:

[0107]

[0108] Therefore, when the L-1 modes are known, the problem of extracting the Lth mode can be expressed as a constrained minimization problem, expressed as:

[0109]

[0110]

[0111] Where, For balance 、 and The parameters of are solved by the Lagrange multiplier method.

[0112] The successive variational mode decomposition results after the moss growth optimization algorithm and its corresponding spectrum are as follows: Figure 4 and Figure 5 As shown in the figure, the successive variational mode decomposition algorithm optimized by the moss growth optimization algorithm is used to decompose the key frequency components of the signal, separate the low-frequency characteristic components from the high-frequency noise components, and help to remove the high-frequency noise components in the subsequent signal reconstruction process.

[0113] The energy distribution of each modal component obtained by decomposition and the original composite impact signal in the time domain is calculated, and the correlation coefficient between the two is calculated. The first four or three modal components with larger correlation coefficients are selected for signal reconstruction to obtain the denoised impact signal. The energy distribution of the signal in the time domain is selected as the basis for screening the modal components. This can effectively retain the transient characteristics and local energy concentration characteristics of the impact signal, avoid losing key transient information, and effectively distinguish the impact signal from the noise component, thereby enhancing the impact positioning capability. The expression is:

[0114]

[0115]

[0116] Where, is the dynamic dual-weight envelope entropy; are the sample points of the composite impact signal and modal components; is the energy distribution correlation coefficient; is the energy distribution of the modal component; is the energy distribution of the original signal; is the covariance; is the variance.

[0117] Signal reconstruction is performed based on the modal components selected based on the obtained correlation coefficients. The signal reconstruction process is as follows: assuming that the original composite impulse signal is decomposed into k modal components, the first M modal components with the largest correlation coefficients are selected, where M is generally 3 or 4, and the screening set S is defined. The reconstruction expression is:

[0118]

[0119]

[0120] Where, It is the correlation index between the modal component energy distribution and the original signal; To reconstruct the signal, is the natural mode component obtained by decomposition;

[0121] like Figure 6 As shown in the figure, the noise amplitude of the reconstructed denoised impulse signal is significantly reduced, achieving the denoising effect;

[0122] Step 4: Use the continuous wavelet transform algorithm to convert the denoised impact signal in step 3 into a two-dimensional time-frequency spectrum, and use the relative position matrix algorithm to convert the composite impact signal in step 1 into a two-dimensional coding image to obtain images of two modes;

[0123] The continuous wavelet transform has the ability to localize time and frequency, and can provide both time and frequency information of the signal. It is suitable for analyzing non-stationary signals such as impact signals. The time-frequency characteristics of the impact signal can be intuitively represented as a two-dimensional time-frequency spectrum, which is conducive to subsequent feature extraction and location analysis. The denoised impact signal in the continuous wavelet transform algorithm is converted into a two-dimensional time-frequency spectrum. The expression is:

[0124]

[0125] Where, is the coefficient of continuous wavelet transform; is the scale parameter; is the translation parameter; is the denoised signal; is the Morlet wavelet basis function; for The conjugate complex of ;

[0126] The relative position matrix algorithm is used to convert the original composite shock signal into a two-dimensional code map, which explicitly represents the spatial distribution characteristics of the shock signal. As a supplementary modal image to the time-frequency spectrum generated by continuous wavelet transform, it provides three-dimensional time-frequency-space information and provides a more comprehensive and accurate feature representation for shock location. The relative position matrix algorithm process includes:

[0127] The original composite shock signal sequence data is z-score normalized to obtain a standard normal distribution, which is expressed as:

[0128]

[0129] Where, is the normalized signal sequence; is the original composite signal sequence; Original composite signal sequence The average value of is the original composite impulse signal sequence The standard deviation of

[0130] Apply the piecewise aggregation approximation method to reduce the normalized time series data from n dimensions to m dimensions by calculating the average value of a piecewise constant, while maintaining the approximate trend of the original series. Select a suitable dimensionality reduction factor k and set the dimensionality reduction factor k to 40 to generate a new smoothed time series. The expression is:

[0131]

[0132] Construct an m×m matrix, calculate the relative position between two timestamps, and convert the processed time series x into a two-dimensional matrix. The expression is:

[0133]

[0134] Apply maximum and minimum normalization to convert M into a grayscale value matrix, the expression is:

[0135]

[0136] Where, is the gray value matrix;

[0137] Step 5: Construct a multimodal sample dataset for training, validation, and prediction of the impact localization model;

[0138] The image data volume of the two-dimensional time-frequency spectrum and the two-dimensional coding image obtained above is expanded through image enhancement operation. The two-dimensional time-frequency spectrum generated by continuous wavelet transform is subjected to image noise addition operation, and four types of noise, namely Gaussian, Poisson, salt and pepper, and multiplicative noise, are added to expand the data set to 5 times the original one, and a total of 1215 two-dimensional time-frequency spectrum images are obtained. The two-dimensional coding image generated by the relative position matrix algorithm is subjected to image transformation operation, and four operations, namely flipping, contrast enhancement, histogram equalization and grayscale, are performed on the image to expand the data set to 5 times the original one, and a total of 1215 two-dimensional coding images are obtained, which finally constitute a sample data set of two modal images. The image sample data set also includes label information corresponding to each image data, that is, the coordinates of the impact position. The data set is divided into a training set and a validation set according to a preset ratio.

[0139] A portion of the original composite impact signal data of random knocking is selected for continuous wavelet transform, and only the two-dimensional time-frequency spectrum generated by continuous wavelet transform is selected as the test data, and this portion of the image is not subjected to image enhancement processing;

[0140] In order to improve the convergence and generalization ability of the model, the labels corresponding to the image data in the training set and the validation set are normalized to the maximum and minimum values. The formula is:

[0141]

[0142] Where, is the normalized coordinate value; is the original coordinate value; is the maximum value of the original data; is the minimum value of the original data;

[0143] Step 6: Construct a ResNet18-NAM-Agent multimodal regression model, fuse the ResNet18 network with the NAM attention mechanism to form a feature extraction network, select the agent attention mechanism as the feature fusion module, train the multimodal regression model with the sample dataset, and obtain a composite laminated structure impact localization model based on the multimodal regression model;

[0144] The ResNet18-NAM-Agent multimodal regression model structure is as follows Figure 7 shown.

[0145] The ResNet18-NAM feature extraction model uses ResNet18 as the backbone network for feature extraction and the NAM attention mechanism as the feature enhancement mechanism. The number of seeds is set to ensure that the datasets of the two modalities are shuffled in the same way. The AdamW optimizer is used to train the model, the learning rate is set to 0.001, and the weight decay coefficient is set to The cosine annealing scheduling algorithm is used to adjust the learning rate. The number of iterations of model training is set to 50, the image training batch is set to 10, and Huber is selected as the loss function. During the model training process, the loss is calculated only on the time-frequency spectrum data set generated by continuous wavelet transform. The mathematical formula of the loss function is:

[0146]

[0147] Where, is the output of the loss function; is the true value; is the predicted value; To adjust the threshold parameter of the Huber loss function expression, it is preferably 3.0.

[0148] The training process of the ResNet18-NAM-Agent model includes: removing the average pooling layer and fully connected layer of the ResNet18 network and constructing a 2D time-frequency image feature extraction branch and a 2D coding graph feature extraction branch respectively;

[0149] Preferably, the ResNet18 network includes a 7×7 convolutional layer with a stride of 2, a 3×3 maximum pooling layer with a stride of 2, four residual layers, an average pooling layer, and a fully connected layer. Each residual layer includes two residual blocks, which are divided into downsampling residual blocks and general residual blocks. The first residual layer consists of a general residual block, and the second, third, and fourth residual layers consist of a downsampling residual block and a general residual block. The general residual block structure is 3×3Conv+BN+ReLU+3×3Conv+BN, and the downsampling residual block structure is 3×3Conv+BN+ReLU+3×3Conv+BN+downsample. The NAM attention module includes a channel attention submodule and a spatial attention submodule.

[0150] The color images of the two modal images in the multimodal image sample dataset are converted to 224×224×3, respectively input into a 7×7 convolutional layer with a stride of 2 to extract image features, and then respectively input into a 3×3 maximum pooling layer with a stride of 2 to reduce the size of the feature map;

[0151] For the feature extraction branch of the 2D time-spectrogram, the average pooling layer and fully connected layer of the ResNet18 model are removed, and the NAM attention module is inserted before the first and second residual layers of the ResNet18 model. The attention mechanism is not added to the last two residual layers.

[0152] The time-frequency feature map after the maximum pooling layer is processed into two residual layers before NAM and two original residual layers in sequence to obtain the CWT feature map extracted by the feature extraction network;

[0153] When entering the NAM attention mechanism, it first enters the channel attention submodule. In this module, the input feature map is normalized by the Batch Normalization layer, and the weight parameters of the Batch Normalization layer are used to calculate the importance of each channel. The normalized weights are multiplied by the input feature map channel by channel to generate a weighted feature map. The weighted feature map is then activated by the Sigmoid function to generate a channel attention weight map, which is multiplied element by element with the original input feature map to obtain the output of the channel attention module.

[0154] Then, the spatial attention submodule is entered. In this module, the input feature map is averaged in the channel dimension to obtain a spatial weight map. After normalization, the spatial weight map is multiplied element-wise with the input feature map to generate a weighted feature map. The weighted feature map is activated by the Sigmoid function to generate a spatial attention weight map, which is multiplied element-wise with the original input feature map to obtain a feature map processed by the NAM module.

[0155] For the feature extraction branch of the 2D relative position encoding map, the average pooling layer and fully connected layer of the ResNet18 model are removed, and the NAM attention module is inserted into the first residual layer, the second residual layer, and the third residual layer of the ResNet18 model before the fourth residual layer;

[0156] The feature map processed by the maximum pooling layer successively enters three residual layers before NAM and one residual layer after NAM to obtain the RPM feature map extracted by the feature extraction network;

[0157] The feature maps of the two modalities obtained by the above process are reshaped into a sequence form to adapt to the input requirements of the agent attention mechanism, and the expression is:

[0158]

[0159]

[0160] Where, is the sequence of spectrograms after the reshaped two-dimensional continuous wavelet transform; is the sequence of the reshaped two-dimensional relative position matrix encoding map; It is the feature map of the two-dimensional time-frequency spectrum after feature extraction; is the feature map of the two-dimensional coding image after feature extraction;

[0161] The proxy attention mechanism introduces a set of additional proxy features into the traditional attention module, which inherits the advantages of softmax and linear attention, enables the model to establish relationships between different modalities and improves feature fusion performance.

[0162] The The sequence is used as query Q and value V input to the agent attention mechanism, As the key K is input into the proxy attention mechanism, the feature matrix of the query Q is average pooled to obtain the proxy feature A, which is expressed as:

[0163]

[0164] The proxy attention mechanism is used to fuse the two modal sequence data, and the expression is:

[0165]

[0166] Where, is the agent feature matrix; is global average pooling; , , , 、 and is the weight parameter, and is the original feature; is the feature sequence obtained after the fusion of the agent attention mechanism; is the softmax function; is the scaling factor; and is the position offset; It is a depth-wise separable convolution operation.

[0167] Averaging the multimodal fusion features in the sequence length dimension to generate a pooled feature vector, inputting the pooled feature vector into a two-dimensional fully connected layer to obtain the two-dimensional coordinates of the impact position;

[0168] Step 7: Use the impact location model to perform impact location on the test set image sample data, and finally determine the impact position.

[0169] The trained ResNet18-NAM-Agent multimodal regression model is used to locate the impact of the image data in the test dataset and obtain the final predicted coordinate value of the impact position.

[0170] Example 2:

[0171] On the basis of the above embodiments, taking into account the problem that traditional envelope entropy only evaluates modal purity through signal energy distribution and does not distinguish the differences in energy attenuation characteristics between impact signals and noise in the time and frequency domains, a dynamic dual-weight envelope entropy is added, and a time domain attenuation factor is introduced to quantify the energy concentration of the signal in the time domain and the sparsity in the frequency domain, respectively, and construct a composite fitness function, so that the moss growth optimization algorithm can more accurately screen out modal components with both low noise and high impact characteristics.

[0172] The dynamic dual-weight envelope entropy formula is defined as:

[0173] Where, For the Dynamic double-weighted envelope entropy of modal components; is the total number of signal sample points; is the frequency domain sparse factor; is the time domain attenuation factor; is the time point of the signal; is the frequency point of the signal; For the The envelope normalized value of the modal component.

[0174] The dynamic double-weighted envelope entropy is used as the fitness function, the population size of the moss growth optimization algorithm is set to 10, and the number of iterations is set to 19. The search range is [500,60000];

[0175] During the wind direction update phase, priority is given to The modal components corresponding to The value is used for population iteration.

[0176] Verification shows that while achieving an average error similar to that of the above-mentioned embodiment, the signal-to-noise ratio of the reconstructed signal is improved from 18.6dB of the traditional method to 24.3dB; in carbon fiber laminate testing, the average absolute positioning error is reduced from 7.2mm to 3.8mm; and the modal aliasing rate is reduced from 12.7% to 5.4%. The results show that this mechanism significantly improves the ability to distinguish between impact features and noise components during signal decomposition by introducing a time-domain attenuation factor and a frequency-domain sparsity factor. The energy distribution of the reconstructed denoised signal in the time-frequency domain is highly consistent with the original impact event, effectively suppressing modal aliasing while preserving the integrity of the impact transient characteristics.

[0177] Example 3:

[0178] On the basis of the aforementioned embodiments, considering that the traditional proxy attention mechanism only fuses multimodal features through linear transformation, ignoring the dynamic correlation between time-frequency features and spatial features in the time and space dimensions, a spatiotemporal collaborative attention mechanism is added, which coordinates the cross-modal interaction strength through the time domain convolution gating unit and the spatial domain dynamic weight to achieve feature complementarity between local details and global distribution.

[0179] The spatiotemporal co-attention formula is defined as:

[0180]

[0181] Where, It is the spatiotemporal coordinated attention output; is the feature vector dimension; is channel-by-channel multiplication; It is the time domain convolution gating; is the spatial dynamic weight.

[0182] After the ResNet18-NAM branch extracts the time-frequency and spatial features, the feature sequence is input into the spatiotemporal collaborative attention module;

[0183] The convolution kernel parameters of the temporal convolution gate unit are initialized to Gaussian distribution, and the MLP hidden layer dimension of the spatial dynamic weight is 256;

[0184] During training, the first three residual layers of the ResNet backbone network are frozen, and only the parameters of the attention module are optimized.

[0185] For example, assuming that the input feature CWT is the spectrogram feature vector , the RPM encoding graph feature vector is ; The parameter setting is, the time domain convolution gate convolution kernel is , the spatial dynamic weight MLP weight matrix is , the bias is , the depth-wise separable convolution weights are , the position offset is , .

[0186] One-dimensional convolution results:

[0187] Sigmoid gate value:

[0188] Global average pooling:

[0189] MLP output:

[0190] Airspace dynamic weight;

[0191] Raw score:

[0192] Adjusted score:

[0193] Attention weighted value:

[0194] DWC Operation:

[0195] Final output:

[0196] The effect of the spatiotemporal collaborative attention mechanism is shown in Table 1.

[0197] Table 1. Summary of the effects of spatiotemporal collaborative attention mechanism

[0198]

[0199] The experimental table shows that while the model parameters increased by only 1.2%, the speed of cross-modal feature interaction increased by 37%. The average Euclidean distance error of the test set decreased from 4.1 mm to 2.5 mm. In a strong noise environment with a signal-to-noise ratio of 10 dB, the error fluctuation range narrowed from ±3.8 mm to ±1.6 mm. The results demonstrate that this mechanism achieves a deep interaction between time-frequency features and spatial distribution features through the coordinated modulation of temporal convolutional gating and spatial dynamic weighting. The model's robustness to complex noise and boundary reflection interference is significantly improved, and the computational efficiency of cross-modal feature fusion is optimized.

[0200] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product embodied on one or more computer-usable non-transitory storage media containing computer-usable program code.

[0201] The present invention can provide computer program instructions to a management platform of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the management platform of the computer or other programmable data processing device produce a device for implementing the system.

[0202] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture including an instruction device that implements the functions of the system.

[0203] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions of the described system.

Claims

1. A multi-modal fusion composite laminate structure low-speed impact positioning method, characterized in that: include: Acquiring impact response signals of multiple strain sensors on the composite laminate structure and combining them into a composite impact signal; The moss growth optimization algorithm is used to optimize the penalty coefficient of successive variational mode decomposition with envelope entropy as the fitness function. The composite shock signal is subjected to successive variational modal decomposition based on the optimal penalty coefficient, and the modal components with the largest energy distribution correlation coefficient are selected to reconstruct the denoised signal. Converting the denoised signal into a time-frequency spectrum diagram, and converting the original composite signal into a two-dimensional coding diagram through a relative position matrix algorithm; Constructing a multimodal sample dataset comprising the spectrogram and coding graph; Train a multimodal regression model and perform shock localization.

2. The method according to claim 1, wherein: The moss growth optimization algorithm includes a wind direction determination mechanism, spore diffusion search, double reproduction search and cryptobiosis mechanism, wherein the wind direction is dynamically adjusted by the positional relationship between the optimal individual and the majority of individuals in the population.

3. The method according to claim 2, wherein: The update formula for the spore diffusion search is: Where, is the wind intensity attenuation factor; is the current evaluation number; is the maximum number of evaluations; for The number of mosses in the total number of moss individuals the proportion of Used to count the number of elements; It is a collection of most moss individuals; is the total number of moss individuals; is the distance spores travel under steady wind conditions; is a constant parameter 2; is the distance spores travel under turbulent wind conditions; For the moss individuals New mosses obtained by spore dissemination; It is a moss individual in the current moss population; For wind direction; is 0.2; 、 and A random number between 0 and 1.

4. The method according to claim 1, wherein: The constraint criteria of the successive variational modal decomposition include minimizing the frequency domain bandwidth of the modal component, the spectrum overlap between the residual signal and the modal component, and the energy of the current mode near the historical modal center frequency.

5. The method according to claim 1, wherein: The energy distribution correlation coefficient is calculated as follows: Where, is the energy distribution correlation coefficient; is the energy distribution of the modal component; is the energy distribution of the original signal; is the covariance; is the variance.

6. The method according to claim 1, wherein: The construction of the multimodal sample dataset includes: Add Gaussian, Poisson, salt and pepper, and multiplicative noise to the time-frequency spectrum, and perform flipping, contrast enhancement, and histogram equalization on the encoded image.

7. The method according to claim 1, wherein: In the multimodal regression model, ResNet18 is fused with the normalized attention mechanism, the proxy attention mechanism is used for multimodal feature fusion, the channel attention calculates the channel importance through the batch normalization weight, and the spatial attention generates the spatial weight through the channel mean; The query and value of the proxy attention mechanism come from the time-frequency graph feature sequence, and the key comes from the encoding graph feature sequence. The fusion formula is: Where, is the agent feature matrix; is global average pooling; , , , 、 and is the weight parameter, and is the original feature; is the feature sequence obtained after the fusion of the agent attention mechanism; is the softmax function; is the scaling factor; and is the position offset; It is a depth-wise separable convolution operation.

8. The method according to claim 1, wherein: It also includes a dynamic dual-weight envelope entropy step, which constructs a composite fitness function by quantifying the energy concentration of the signal in the time domain and the sparsity in the frequency domain. The dual-weight envelope entropy formula is defined as: Where, For the Dynamic double-weighted envelope entropy of modal components; is the total number of signal sample points; is the frequency domain sparse factor; is the time domain attenuation factor; is the time point of the signal; is the frequency point of the signal; For the The envelope normalized value of the modal component.

9. The method according to claim 7, wherein: It also includes a spatiotemporal collaborative attention step, which coordinates the cross-modal interaction strength through the temporal convolutional gating unit and the spatial dynamic weight. The spatiotemporal collaborative attention formula is defined as: Where, It is the spatiotemporal coordinated attention output; is the feature vector dimension; is channel-by-channel multiplication; It is the time domain convolution gating; is the spatial dynamic weight.

10. A computer-readable storage medium storing a computer program, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is executed.

Citation Information

Patent Citations

  • Turbine shaft double-rotor inter-shaft rub-impact identification method and device based on VMD and double attention mechanism TCN

    CN119128687A

  • Health monitoring system and monitoring method for wind turbine blades

    WO2024255027A1