Multi-modal aliasing sound signal separation method based on lightweight convolutional neural network

By processing acoustic emission signals of composite material damage using a lightweight convolutional neural network, the problem of separating multimodal aliased acoustic signals is solved, achieving fine separation and identification in high-noise environments. This is suitable for real-time edge applications and generates reliable damage pattern labels.

CN120895050AInactive Publication Date: 2025-11-04SPECIAL EQUIP SAFETY SUPERVISION INSPECTION INST OF JIANGSU PROVINCE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511199099.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies suffer from low efficiency, high misjudgment rate, and strong reliance on human experience when processing acoustic emission signals of composite material damage. Furthermore, the accuracy and reliability of identification decrease in high-noise environments, especially in multimodal aliased acoustic signals where it is difficult to effectively extract and identify weak damage features.

Method used

A lightweight convolutional neural network is used to separate multimodal aliased acoustic signals. The time-frequency complex spectrum is obtained by short-time Fourier transform, mapped to Mel spectrogram and preprocessed. Feature extraction is performed using the Sandglass module and ECA attention mechanism configured by EfficientNet-B0. Combined with DropBlock regularization, time-frequency soft mask separation and category label generation are achieved.

Benefits of technology

It achieves fine separation under aliasing conditions, reduces computational load and parameter count, is suitable for real-time edge applications, has low latency and low power consumption characteristics, can accurately identify composite material damage modes in high-noise environments, and generate reliable category labels for alarms and diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895050A_ABST
    Figure CN120895050A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal aliasing sound signal separation method based on a lightweight convolutional neural network, relates to the technical field of signal processing, and aims to remove noise and perform normalization processing by preprocessing collected multi-modal aliasing sound signals. And then extracting features by using a Mel spectrogram, and converting the sound signal into a frequency spectrum representation conforming to auditory perception of human ears. Then, the extracted features are trained and learned through a constructed lightweight convolutional neural network model OfficientNet-B0, and a trained model weight is obtained; and finally, separating and identifying the new multi-modal aliasing sound signals by using the model, thereby realizing accurate classification of different types of damage signals, and effectively solving the problem that the multi-modal aliasing sound signals are difficult to accurately separate in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of signal processing, in particular to a multi-modal mixed acoustic signal separation method based on a lightweight convolutional neural network. BACKGROUND

[0002] Fiber-reinforced plastics (FRP) shows a growing trend in applications such as automotive, aircraft components, industrial engineering equipment and manufacturing industries due to its relatively high strength / weight ratio, special material properties of changing stacking order and fiber orientation. The relative defect of FRP composite material is the lack of ductile behavior, and its sensitivity to damage can cause premature failure of the structure under static and fatigue loads. Since structural damage of FRP composite materials often occurs internally, there is no corresponding warning stage before this, so it is of great significance to monitor and analyze the damage evolution process of the composite material and evaluate it. With the continuous development of signal processing technology and the progress of sensor technology, acoustic emission (AE) technology in the field of composite material characterization has promoted the development of detection methods and technologies. Acoustic emission monitoring technology can realize the qualitative description of damage behaviors such as deformation and fracture of composite structures, but due to the many damage factors of composite structures, different damage mechanisms have different effects on damage, and acoustic emission signals under multiple damage modes are mixed and interfered with each other, with various parameter distributions and disorder. These acoustic signals often contain information of multiple damage modes, forming multi-modal mixed acoustic signals. Traditional manual analysis methods have low efficiency, high misjudgment rate, and strong dependence on human experience when processing such signals. The conventional deep learning network model usually needs a large amount of labeled data for training, and it is difficult and time-consuming to collect a large amount of damage data containing label information in complex industrial production environments, and it is difficult to effectively extract and identify weak damage features. Especially in a high-noise environment, traditional models are sensitive to noise, resulting in a decrease in recognition accuracy and reliability, which cannot meet the actual needs. SUMMARY

[0003] Based on the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a multi-modal mixed acoustic signal separation method based on a lightweight convolutional neural network to solve the above technical problems.

[0004] To achieve the above-mentioned purpose, the present application provides the following technical scheme: a multi-modal mixed acoustic signal separation method based on a lightweight convolutional neural network, comprising:

[0005] Obtaining the original waveform data of the multi-modal mixed acoustic signal, performing short-time Fourier transform on the original waveform data to obtain a time-frequency complex spectrum;

[0006] Map the time-frequency complex spectrum to a mel-frequency spectrogram, and perform image preprocessing on the mel-frequency spectrogram to obtain a standardized time-frequency graph;

[0007] Input the standardized time-frequency graph into a pre-trained lightweight convolutional neural network to obtain a posterior probability graph of each target category with the same size as the input time-frequency graph, the lightweight convolutional neural network adopts the stage configuration of EfficientNet-B0 and takes the Sandglass module as the core unit, the Sandglass module includes channel expansion, depth separable convolution and channel projection, the activation function is Mish, and ECA channel attention is introduced at each stage; when the residual connection condition is met, DropBlock regularization is applied at the end of the transformation branch;

[0008] Based on the posterior probability graph, the mel-frequency spectrogram corresponding to the same time-frequency coordinate is separated by time-frequency soft mask, to obtain separated spectral components and category labels divided by categories, and the multi-modal mixed sound signal is separated according to the category label.

[0009] The application further provides that the multi-modal mixed sound signal includes sound signals in three damage mode states of fiber breakage, matrix cracking and delamination.

[0010] The application further provides that the image preprocessing of the mel-frequency spectrogram includes:

[0011] Size adaptation cropping and sub-block generation: according to the high-frequency wide distribution of fiber breakage signals, the low-frequency aggregation of matrix cracking, and the frequency band gradual change characteristics of delamination signals, the basic size of 224x112 for fiber breakage and 224x224 for matrix cracking and delamination is cropped, and a sliding window mechanism with a 56-pixel step is used to generate 112x112 sub-image blocks from the basic image;

[0012] Directional data enhancement: only horizontal flip operation is implemented to avoid destroying the time flow characteristics of the sound signal, and 0.8-1.2 times random stretching is performed on the full frequency band of 500-8000Hz to simulate the frequency broadening effect of different damage evolution stages;

[0013] Standardization calibration: based on the global statistical characteristics of the data set, the mean and standard deviation are set, and the tensor image is standardized to ensure the consistency of the feature distribution of different damage samples.

[0014] The application further provides that the deep learning network model is constructed, including:

[0015] When constructing a deep learning network, a plurality of Sandglass modules are stacked to form a backbone structure, each module fuses extended convolution, grouped convolution, ECA attention mechanism and shortcut connection, the front part of the module expands the input channel to t times the original channel number through 1x1 convolution, t belongs to {2, 4}, and is followed by a Mish activation function to improve the nonlinear ability of feature expression, the mathematical expression is: Mish(x) = x*tanh(ln(1+ex)), x is the element value in the feature map;

[0016] In order to enhance the correlation between channels, the module introduces a lightweight ECA attention mechanism to locally weight the features in the channel dimension using a one-dimensional convolution kernel, the calculation process is: a = sigma(Conv1D([GAP(F1)])), wherein a is the channel attention weight vector generated by ECA, sigma represents the Sigmoid activation function, F1 is the input feature map, Conv1D is the channel one-dimensional convolution, and GAP(F1) is the global average pooling of the input feature map F1;

[0017] When the input and output channel numbers are the same and the convolution step is 1, the module enables the shortcut connection structure, adds the original input and the convolution transformed features, and introduces the DropBlock regularization term, and the output form is: Fout = DropBlock(F2) + X, wherein Fout is the output feature map, X is the input feature, and F2 is the feature map processed by ECA attention and convolution.

[0018] The application further provides that the connection mode between the plurality of Sandglass modules is the connection mode between layers in the deep learning network.

[0019] The application further provides that the lightweight convolutional neural network is provided with a time-frequency dense output head for converting the output feature map into a posterior probability map of each target category with the same time step and frequency band number as the input time-frequency map; wherein the time-frequency dense output head comprises:

[0020] A category mapping layer is used to map the backbone features to the confidence of each target category at each time-frequency position;

[0021] A resolution alignment module is used to restore or maintain the features to the time-frequency coordinate grid of the input time-frequency map through upsampling, transpose convolution and / or replacing the downsampling with a dilated convolution in the case of downsampling in the backbone; Figure One A normalization processing is used to normalize the confidence of each target category at each time-frequency position to obtain the posterior probability map of each target category.

[0022] The application is further configured that, when the lightweight convolutional neural network is trained and verified, the data used for training and verification is divided by group hierarchical cross-validation, the original sample batch is taken as a group unit for division at the group level, the data is divided into K subsets according to the sample proportion of the three types of damage signals of fiber fracture, matrix cracking and delamination 10:1:10, the group leakage potential is calculated according to the grouping result, the maximum value in each group leakage potential is taken as the overall leakage potential and is limited to be less than or equal to a preset threshold, and when the overall leakage potential is greater than the preset threshold, re-grouping is performed.

[0023] The application is further configured that, at each time-frequency position, the posteriori confidence of each target category is normalized position by position to obtain a weight mask with a total sum limit at the position.

[0024] According to the weight mask, the energy is proportionally distributed at the same time-frequency position to obtain a separated spectral component divided by category.

[0025] The category label is calculated based on the separated spectral component.

[0026] The application is further configured that the calculation logic of the category label comprises:

[0027] The separated spectral components of each target category are energy-aggregated in the frequency band dimension at each time frame to obtain the category energy of the frame.

[0028] The candidate category of the frame is determined according to the proportion of the energy of each target category in the total energy in the same time frame.

[0029] When the highest energy proportion is greater than or equal to a preset threshold, the candidate category is determined as the category label of the frame, and when the highest energy proportion is less than the preset threshold, it is recorded as undetermined.

[0030] The application is further configured that the energy proportional distribution is performed in the linear energy domain; when the input is a log-scaled mel spectrogram, the log-scaled mel spectrogram is restored to the linear energy domain before the time-frequency soft mask separation is performed.

[0031] The application provides a multi-modal mixed acoustic signal separation method based on a lightweight convolutional neural network, obtains original waveform data of multi-modal mixed acoustic signals, and performs short-time Fourier transform on the original waveform data to obtain a time-frequency complex spectrum; the time-frequency complex spectrum is mapped into a mel spectrum graph, the mel spectrum graph is pre-processed into a standardized time-frequency graph; the standardized time-frequency graph is input into a pre-trained lightweight convolutional neural network to obtain a posterior probability graph of each target category with the same size as the input time-frequency graph, the lightweight convolutional neural network adopts the stage configuration of EfficientNet-B0 and takes the Sandglass module as the core unit, the Sandglass module includes channel expansion, depth separable convolution and channel projection, the activation function is Mish, and ECA channel attention is introduced at each stage; when the residual connection condition is met, DropBlock regularization is applied at the end of the transform branch; based on the posterior probability graph, the mel spectrum graph corresponding to the same time-frequency coordinate is separated by time-frequency soft mask separation to obtain separated spectral components and category labels classified by categories, and multi-modal mixed acoustic signal separation is performed according to the category labels, and the beneficial effects include:

[0032] 1. Fine separation under the condition of aliasing: the category posterior probability graph with the same resolution as the input is obtained through the time-frequency dense output head, and soft mask separation is implemented on the same time-frequency coordinate, so that strong overlapping acoustic signals such as fiber breakage, matrix cracking and delamination can be decoupled at the spectral level, and information loss caused by whole segment discrimination is avoided;

[0033] 2. Lightweight and easy to deploy: the combination of EfficientNet-B0 stage configuration with Sandglass as the core, depth separable convolution, ECA attention and DropBlock significantly reduces the parameter quantity and operation quantity while ensuring the separation and recognition accuracy, and is suitable for end-side / on-line real-time application, and meets the needs of low latency and low power consumption scenes;

[0034] 3. Closed-loop output from separation to label: after obtaining the separated spectral components, the category labels at the frame level, segment level or whole sample level are directly generated based on the energy proportion method, forming a closed-loop process of "posterior probability to soft mask to separated spectral components to label", which can be directly used for alarm, positioning and subsequent diagnosis;

[0035] 4. Anti-leakage and credible evaluation: the grouping hierarchical cross-validation is adopted with sample batches as units and the maximum grouping leakage potential is used as a constraint to ensure that any grouping belongs to only one subset, effectively avoiding the leakage of homologous information between training-validation / testing, and the evaluation results are more real and reliable.

[0036] The above description is only a summary of the technical solutions of the present application. In order to make the technical means of the present application more clearly understood and implemented according to the content of the description, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0038] Figure 1 The flowchart of the multi-modal mixed acoustic signal separation method based on a light convolutional neural network shown for an exemplary embodiment of the present application;

[0039] Figure 2 The fiber fracture damage mode mel spectrum shown for an exemplary embodiment of the present application;

[0040] Figure 3 The matrix cracking damage mode mel spectrum shown for an exemplary embodiment of the present application;

[0041] Figure 4 The delamination damage mode mel spectrum shown for an exemplary embodiment of the present application;

[0042] Figure 5 The training accuracy curve in different stages shown for an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0043] The embodiments of the present application will be described below with reference to the drawings and preferred embodiments, and those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in the present specification. The present application can also be implemented or applied by different specific embodiments, and the details in the present specification can be modified or changed based on different views and applications without departing from the spirit of the present application. It should be understood that the preferred embodiments are only for illustration of the present application, and are not intended to limit the protection scope of the present application.

[0044] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner, and only show the components related to the present application in the diagrams, not the number, shape and size of the components when actually implemented. The actual implementation of each component may be a random change, and the component layout pattern may be more complex.

[0045] In the following description, numerous specific details are discussed in order to provide a thorough understanding of embodiments of the present application. However, it will be apparent to one skilled in the art that the present application can be practiced without these specific details. In other instances, well-known structures and devices are not described in detail in order to avoid obscuring embodiments of the present application.

[0046] A multi-modal mixed acoustic signal separation method based on a lightweight convolutional neural network, as shown in Figure 1 includes:

[0047] Obtain the original waveform data of the multi-modal mixed acoustic signal, and perform short-time Fourier transform on the original waveform data to obtain a time-frequency complex spectrum;

[0048] Map the time-frequency complex spectrum to a mel spectrum graph, and perform image preprocessing on the mel spectrum graph to obtain a standardized time-frequency graph;

[0049] Input the standardized time-frequency graph into a pre-trained lightweight convolutional neural network to obtain a posterior probability graph of each target category with the same size as the input time-frequency graph, the lightweight convolutional neural network adopts the stage configuration of EfficientNet-B0 and takes the Sandglass module as the core unit, the Sandglass module includes channel expansion, depth separable convolution and channel projection, the activation function is Mish, and ECA channel attention is introduced at each stage; When the residual connection condition is met, DropBlock regularization is applied at the end of the transformation branch;

[0050] Based on the posterior probability graph, the mel spectrum graph corresponding to the same time-frequency coordinate is separated by time-frequency soft mask separation to obtain separated spectral components and category labels divided by category, and multi-modal mixed acoustic signal separation is performed according to the category labels.

[0051] Specifically, the acquisition of the acoustic emission dataset: the acoustic emission dataset A collected in the typical damage test is divided into a training set, a verification set and a test set according to a ratio of 6:2:2, three kinds of damage signals of fiber fracture, matrix cracking and delamination are collected, and the data constitutes the following table 1-1. The acoustic emission dataset B collected in the composite gas cylinder water pressure test is a test set, and the acoustic emission signals of 8 pressure maintaining stages are collected, and the distribution of the acoustic emission signals collected in each stage is as shown in the following table 1-2. The model is used to classify the three kinds of damage signals in each stage.

[0052]

[0053] Table 1-1 Composition of dataset A

[0054]

[0055] Table 1-2 Composition of dataset B

[0056] The mel-spectrogram is used as a feature extraction method, the frequency spectrum information is extracted by a short-time Fourier transform, and the frequency is mapped to a non-linear mel scale by using a mel filter bank, and then the energy is logarithmized, the frequency axis of the mel-spectrogram adopts the mel scale, and the conversion formula is: Wherein, M(f) is the value of linear frequency f mapped to the mel scale, and f is a linear frequency, that is, an actual frequency on the STFT spectrum;

[0057] The application is further provided as, the multimodal aliasing sound signal includes the sound signal in three damage mode states of fiber fracture, matrix cracking and delamination. The mel-spectrogram is converted to three kinds of damage signals, and the obtained mel-spectrogram is as shown in Figures 2-4 As can be seen from the visual features of the mel-spectrogram, the fiber fracture (b) generally shows higher overall energy, brighter color, and relatively uniform energy distribution, covers a wider frequency range, and has stronger time persistence. The matrix cracking (a) presents lower energy, darker color, and energy concentrated in the low-frequency region, and may present obvious horizontal stripe structure. The energy level of delamination (c) is between the two, and the energy is concentrated in the low-frequency region, and may also present stripe structure, but may be accompanied by energy attenuation phenomenon. These visual differences reflect the different characteristics of acoustic emission signals in energy, frequency and time under different fracture modes. The above only represents the typical mel-spectrogram corresponding to different damages, and more specific differences need to be put into the network for identification for each damage.

[0058] The application is further provided as, the mel-spectrogram is image preprocessed, including:

[0059] Size adaptation cropping and sub-block generation: according to the high-frequency wide-range distribution of fiber fracture signal, the low-frequency aggregation of matrix cracking, and the frequency band gradual change characteristics of delamination signal, the basic size of 224*112 for fiber fracture, 224*224 for matrix cracking and delamination is cropped, and a sliding window mechanism with a 56-pixel step is used to generate 112*112 sub-image blocks from the basic image;

[0060] Directional data enhancement: only horizontal flip operation is implemented to avoid destroying the time flow characteristics of the sound signal, and 0.8-1.2 times random stretching is performed on the full frequency band of 500-8000Hz to simulate the frequency broadening effect of different damage evolution stages;

[0061] Standardization calibration: based on the global statistical characteristics of the data set, the mean and standard deviation are set, and the tensor image is standardized to ensure the consistency of the feature distribution of different damage samples.

[0062] EfficientNet-B0 is a lightweight convolutional neural network, which adopts a compound scaling strategy to balance the relationship between model width, depth and resolution. Its core architecture is composed of Sandglass modules and ECA (Efficient Channel Attention) attention mechanism. EfficientNet-B0 includes an initial convolutional layer, 7 Sandglass stages and a classification head. The input resolution is 224x224, which is down-sampled to 7x7 feature maps step by step, and then global average pooling and fully connected layers are used to complete the classification task. The parameter configuration of each stage is shown in Tables 1-3.

[0063]

[0064]

[0065] Table 1-3 EfficientNet-B0 parameter configuration table

[0066] Among them, the ECA ratio represents the compression layer channel reduction multiple (such as 4 represents the intermediate layer channel number is 1 / 4 of the input channel); the expansion ratio (t) controls the Sandglass module expansion layer channel number (input channelxt).

[0067] The application further provides a deep learning network model, which comprises:

[0068] When constructing the deep learning network, a plurality of Sandglass modules are stacked to form a backbone structure, each module integrates an expanded convolution, a grouped convolution, an ECA attention mechanism and a shortcut connection, the input channel is expanded to t times of the original channel number by 1x1 convolution in the front part of the module, t∈{2,4}, followed by a Mish activation function to improve the nonlinear ability of feature expression, the mathematical expression is: Mish(x)=x·tanh(ln(1+ex)), x is the element value in the feature map; specifically, the core construction unit of EfficientNet-B0 is the Sandglass module, which integrates an expanded convolution, a depth separable convolution, an ECA attention mechanism and a residual connection to realize efficient feature extraction and channel dynamic calibration. Each Sandglass module first expands the input channel to t times of the input channel by 1x1 convolution, and uses a Mish activation function to enhance the nonlinear expression ability;

[0069] To enhance the inter-channel feature correlation, the module introduces a lightweight ECA attention mechanism to locally weight the features in the channel dimension using a one-dimensional convolution kernel, without compression operation to obtain channel-dependent response, thereby reducing parameter overhead and improving real-time performance, the calculation process is: a = sigma (Conv1D ([GAP (F1)]) ), wherein, a is the channel attention weight vector generated by ECA, sigma represents the Sigmoid activation function, F1 is the input feature map, Conv1D is the one-dimensional convolution in the channel dimension, and GAP (F1) is the global average pooling operation on the input feature map F1;

[0070] When the number of input and output channels is the same and the convolution step is 1, the module enables a shortcut connection structure, adds the original input and the feature after convolution transformation, and introduces a DropBlock regularization term, and the output form is: Fout = DropBlock (F2) + X, wherein, Fout is the output feature map, X is the input feature, and F2 is the feature map after ECA attention and convolution processing.

[0071] The overall model adopts a compound scaling strategy to uniformly regulate the width alpha, depth beta and input gamma of the network, and the three satisfy the calculation constraint condition alpha * beta 2 ·gamma 2 ≈2, so as to ensure the linear growth of resource consumption when expanding the model capacity, and avoid the concentration of calculation burden in a single dimension. EfficientNet-B0 is used as the basic structure with alpha = 1.0, beta = 1.0 and gamma = 1.0, so as to balance the model performance and calculation efficiency;

[0072] The channel expansion factor t of the Sandglass module at each stage is configured differently: t = 2 is used in the 2-layer stage to suppress redundant calculation, and t = 4 is used in the deep layer stage to enhance the expression dimension; the convolution kernel adopts a layer-by-layer progressive structure, wherein the deep layer module (such as stage 3, 5 and 6) introduces a 5*5 convolution to enhance the global receptive field, and the shallow layer module (such as stage 1, 2, 4 and 7) selects a 3*3 convolution to reduce the calculation complexity. All modules embed ECA attention mechanism to maintain the feature channel weight modeling ability under the lightweight structure; the regularization part uses DropBlock instead of the conventional Dropout, and is applied to the classification head to double inhibit the overfitting phenomenon.

[0073] The application further sets that the connection mode between the multiple groups of Sandglass modules is the connection mode between layers in the deep learning network.

[0074] The application is further configured that the lightweight convolutional neural network is provided with a time-frequency dense output head for converting the output feature map into a target class posterior probability map with the same time step and frequency band number as the input time-frequency map; wherein the time-frequency dense output head comprises:

[0075] a class mapping layer for mapping the backbone feature into a confidence of each target class at each time-frequency position;

[0076] a resolution alignment module for restoring or maintaining the feature to the time-frequency coordinate grid of the input time-frequency map by upsampling, transpose convolution and / or replacing downsampling with a dilated convolution in the case of downsampling of the backbone; and a normalization processing for normalizing the confidence of each target class at each time-frequency position to obtain a posterior probability map of each target class. Figure One

[0077] Specifically, the residual output feature map Fout of the last layer Sandglass module of each stage of the backbone is taken as the input of the output head: after 1x1 convolution to unify the channels of Fout_8, Fout_4 and Fout_2 with resolutions of T / 8, T / 4 and T / 2 respectively, step-by-step upsampling and side branch fusion (each level is followed by a 3x3 depth separable convolution and ECA) are adopted to finally restore to TxM resolution, map to K class channels and normalize at each time-frequency position to obtain a posterior probability map P k (t,m)(k=1,…,K), wherein T is the number of time frames, M is the number of mel frequency bands determined by the number of mel filter banks, represents how many frequency bands of mel scale on the vertical axis, t is the time frame index, m is the mel frequency band index, k is the class index, and K is the total number of classes, which is 3 in the application.

[0078] The application is further configured that when training and verifying the lightweight convolutional neural network, the data division for training and verification adopts group hierarchical cross-validation, the original sample batch is taken as the grouping unit for division at the group level, the data is divided into N subsets according to the sample proportion of 10:1:10 of the three types of damage signals of fiber breakage, matrix cracking and delamination, the group leakage potential is calculated according to the grouping result, the maximum value in each group leakage potential is taken as the overall leakage potential and limited to be less than or equal to a preset threshold, and when it is greater than the preset threshold, re-grouping is performed; specifically, for each group g: let the number of samples of this group falling into the training set train / verification set val / test set test be n tr ,n va ,n te , the group size n g , and the calculation logic of the group leakage potential is: ​Preferably, the preset threshold is 0, so that any group belongs to only one subset, ensuring that no batch is split into multiple sets, thereby achieving a leak-free, reproducible and reliable training / verification / testing partitioning process.

[0079] like Figure 5 As shown, the model was trained on a computer equipped with an Nvidia 4060Ti GPU. The model was implemented in Python 3.11 and PyTorch. The training hyperparameters are shown in Table 1-4. After loading the initial weights of the network using the pre-trained model, the RAdam optimization algorithm (Rectified Adam) was used for training optimization. The initial learning rate was 0.001, and a weight decay control strategy was combined. When the validation loss did not decrease for three consecutive rounds, the regularization coefficient was increased from 0.01 to 0.03 to enhance convergence. A variant of Focal Loss was selected as the loss function. A focusing factor of γ=2 and an additional weight amplification term were introduced for the matrix crack samples to improve the effect of class imbalance. The learning rate was updated using an exponential annealing decay method. After every 10 rounds of training, the current learning rate was multiplied by 0.8. The step size was gradually reduced while ensuring stable model training, thereby improving the overall training efficiency and final recognition accuracy.

[0080] Hyperparameters Value Optimizer RAdam L2regularization 0.01 LearningRate 0.01 BatchSize 64 MaxEpochs 100 ImageInputSize 224×224 LossFunction CrossEntropyLoss

[0081] Table 1-4 Model Training Parameters

[0082] The present invention is further configured such that, at each time-frequency position, the posterior confidence of each target category is normalized position by position to obtain a weight mask with a limited sum at that position; specifically, the posterior probabilities of each category at the same time-frequency point are compressed into a set of weights that add up to 1, which are used to allocate energy proportionally. The calculation logic of the weight mask is as follows: P k (t,m) represents the posterior probability of the k-th class at time frame t and Mel band m. Similarly, P j (t,m) represents the posterior probability of the j-th class at time frame t and Mel band m; M k (t,m) is the weight mask for the k-th category at time frame t and Mel band m;

[0083] According to the weight mask, energy is proportionally allocated at the same time-frequency position to obtain separated spectral components classified by category; specifically, the input energy is allocated to each category according to the weight mask ratio, and the calculation logic of the separated spectral components is as follows: in, For the separated spectral components, S lin (t,m) represents the Mel spectrum energy in the linear energy domain;

[0084] The class label is calculated based on the separated spectral components. The application is further configured that the calculation logic of the class label comprises:

[0085] The separated spectral components of each target class in each time frame are aggregated in the energy band dimension to obtain the class energy of the frame. Specifically, the aggregated band energy of each time frame t is: E k [t] is the aggregated energy of the kth class in the tth frame; E Σ [t] is the total energy of the tth frame;

[0086] The candidate class of the frame is determined according to the proportion of the energy of each target class to the total energy in the same time frame. Specifically, the energy proportion of the kth class in the tth frame is calculated, and the maximum one is selected as the candidate, r k [t] is the energy proportion of the kth class in the tth frame, and ε is a constant to prevent zero, which is 10 -8 ; label[t] is the class label of the tth frame, and arg represents the independent variable that makes the function obtain the maximum (or minimum) value;

[0087] When the highest energy proportion is greater than or equal to the preset threshold, the candidate class is determined as the class label of the frame, and when the highest energy proportion is less than the preset threshold, it is recorded as undetermined. Specifically, only when is greater than or equal to 0.5, it is accepted, otherwise it is set as undetermined.

[0088] The application is further configured that the energy is proportionally distributed in the linear energy domain. When the input is a log-scaled mel spectrogram, the log-scaled mel spectrogram is restored to the linear energy domain before performing the time-frequency soft mask separation. Specifically, since the time-frequency complex spectrum is obtained by performing short-time Fourier transform on the original waveform data, and the frequency is mapped to the nonlinear mel scale using the mel filter bank, and the energy is logarithmized, the energy is proportionally distributed in the linear energy domain, so it is necessary to restore the log-scaled mel spectrogram to the linear energy domain. That is, the inverse operation of the logarithmic logic is performed.

[0089] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for separating multimodal aliased acoustic signals based on a lightweight convolutional neural network, characterized in that, include: The original waveform data of the multimodal aliased acoustic signal is acquired, and the time-frequency complex spectrum is obtained by performing a short-time Fourier transform on the original waveform data. The time-frequency complex spectrum is mapped to a Mel spectrogram, and the Mel spectrogram is preprocessed to obtain a normalized time-frequency spectrum. The standardized time-frequency image is input into a pre-trained lightweight convolutional neural network to obtain posterior probability maps of each target category with the same size as the input time-frequency image. The lightweight convolutional neural network adopts the stage configuration of EfficientNet-B0 and uses the Sandglass module as the core unit. The Sandglass module includes channel expansion, depthwise separable convolution, and channel projection, with the Mish activation function. ECA channel attention is introduced in each stage. DropBlock regularization is applied at the end of the transform branch when the residual connection condition is met. Based on the posterior probability map, time-frequency soft mask separation is performed on the Mel spectrum map of the corresponding time-frequency coordinates to obtain the separated spectral components and category labels according to the categories. Multimodal aliasing acoustic signals are then separated according to the category labels.

2. The method for separating multimodal aliased acoustic signals based on a lightweight convolutional neural network according to claim 1, characterized in that, Multimodal aliasing acoustic signals include acoustic signals under three damage modes: fiber fracture, matrix cracking, and delamination.

3. The method for separating multimodal aliased acoustic signals based on a lightweight convolutional neural network according to claim 2, characterized in that, Image preprocessing of the Mel spectrogram includes: Size-adaptive cropping and sub-block generation: Based on the high-frequency wide-range distribution of fiber fracture signals, the low-frequency aggregation of matrix cracking, and the frequency band gradation characteristics of delamination signals, the image is specifically cropped to a base size of 224×112 for fiber fracture and 224×224 for matrix cracking and delamination. A sliding window mechanism with a step size of 56 pixels is used to generate 112×112 sub-image blocks from the base image. Directional data augmentation: Only horizontal flipping is performed to avoid disrupting the temporal characteristics of the acoustic signal, while random stretching of 0.8-1.2 times is applied to the entire 500-8000Hz frequency band to simulate the frequency broadening effect at different stages of damage evolution; Standardization calibration: Based on the global statistical characteristics of the dataset, the mean and standard deviation are set to standardize the tensor image to ensure the consistency of feature distribution of different damage samples.

4. The method for separating multimodal aliased acoustic signals based on a lightweight convolutional neural network according to claim 1, characterized in that, Building deep learning network models includes: When constructing a deep learning network, multiple sets of Sandglass modules are stacked to form the backbone structure. Each module integrates extended convolution, grouped convolution, ECA attention mechanism and shortcut connection. The front end of the module expands the input channels to t times the original number of channels through 1×1 convolution, where t∈{2,4}. Then, the Mish activation function is applied to improve the non-linearity of feature representation. The mathematical expression is: Mish(x)=x·tanh(ln(1+ex)), where x is the element-wise value in the feature map. To enhance the correlation between channel features, the module introduces a lightweight ECA attention mechanism that uses a one-dimensional convolution kernel to locally weight features in the channel dimension. The calculation process is as follows: a = σ(Conv1D([GAP(F1)])), where a is the channel attention weight vector generated by ECA, σ represents the Sigmoid activation function, F1 is the input feature map, Conv1D is the one-dimensional convolution in the channel dimension, and GAP(F1) is the global average pooling of the input feature map F1. When the number of input and output channels is the same and the convolution stride is 1, the module enables the shortcut connection structure, adds the original input and the features after convolution transformation, and introduces the DropBlock regularization term. The output form is: Fout = DropBlock(F2) + X, where Fout is the output feature map, X is the input feature, and F2 is the feature map after ECA attention and convolution processing.

5. The method for separating multimodal aliased acoustic signals based on a lightweight convolutional neural network according to claim 4, characterized in that, The connection method between multiple Sandglass modules is the same as the connection method between layers in a deep learning network.

6. The method for separating multimodal aliased acoustic signals based on a lightweight convolutional neural network according to claim 4, characterized in that, The lightweight convolutional neural network uses a time-frequency dense output header to convert the output feature map into posterior probability maps of each target category at the same time step and frequency band as the input time-frequency map; the time-frequency dense output header includes: The category mapping layer is used to map the backbone features to the confidence scores of each target category at each time frequency location; The resolution alignment module is used to restore or maintain the features to a time-frequency coordinate grid consistent with the input time-frequency map by upsampling, transposed convolution and / or replacing downsampling with dilated convolution when downsampling exists in the backbone; and the normalization process is used to normalize the confidence of each target category at each time-frequency position to obtain the posterior probability map of each target category.

7. The method for separating multimodal aliased acoustic signals based on a lightweight convolutional neural network according to claim 1, characterized in that, When training and validating lightweight convolutional neural networks, the data used for training and validation is divided into grouped hierarchical cross-validation. The original sample batches are used as grouping units to divide the data at the group level. The data is divided into K subsets according to the sample ratio of three types of damage signals: fiber breakage, matrix cracking, and delamination, which is 10:1:

10. The leakage potential of each group is calculated based on the grouping results. The maximum value of the leakage potential in each group is used as the overall leakage potential and is limited to be less than or equal to a preset threshold. When it is greater than the preset threshold, the data is regrouped.

8. The method for separating multimodal aliased acoustic signals based on a lightweight convolutional neural network according to claim 1, characterized in that, At each time-frequency position, the posterior confidence of each target category is normalized position by position to obtain a weight mask with limited sum at that position. According to the weighted mask, the energy is proportionally distributed at the same time-frequency position to obtain the separated spectral components divided by category; Category labels are calculated based on the separated spectral components.

9. The method for separating multimodal aliased acoustic signals based on a lightweight convolutional neural network according to claim 8, characterized in that, The calculation logic for category labels includes: In each time frame, the energy of the separated spectral components of each target category is aggregated in the frequency band dimension to obtain the category energy of that frame; Within the same time frame, candidate categories are determined based on the proportion of energy of each target category to the total energy. When the highest energy percentage is greater than or equal to a preset threshold, the candidate category is determined as the category label for that frame; when the highest energy percentage is less than the preset threshold, it is recorded as undetermined.

10. The method for separating multimodal aliased acoustic signals based on a lightweight convolutional neural network according to claim 8, characterized in that, Energy is distributed proportionally in the linear energy domain; when the input is a logarithmically scaled Mel spectrum, the logarithmically scaled Mel spectrum is restored to the linear energy domain before time-frequency soft mask separation is performed.

Citation Information

Patent Citations

  • Transmission device mixed fault separation method based on memory selection mechanism

    CN114360576A

  • Single-channel blind source separation method based on improved Conv-TasNet network

    CN119202639A

  • Speech separation method and apparatus, and storage medium

    WO2023020500A1