Coal gangue acoustic signal separation device and method based on deep time-frequency feature fusion
Through a method based on deep time-frequency feature fusion, combined with convolutional neural network, bidirectional long and short time memory network and multi-head attention mechanism, the problem of coal gangue sound signal separation under the influence of complex background noise in coal mining is solved, and efficient and accurate signal separation effect is achieved.
Patent Information
- Application Number
- CN202510200761.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-27
AI Technical Summary
During coal mining, the identification and separation of coal and gangue are affected by complex background noise, and traditional methods have limitations in noise adaptability, signal distortion and real-time processing capabilities.
The coal gangue acoustic signal separation device and method based on deep time-frequency feature fusion is adopted. Through the time domain feature extraction module, the frequency domain feature extraction module and the feature fusion and source separation module, combined with the convolutional neural network, a bidirectional long and short-term memory network and a multi-head attention mechanism, the extraction of deep features and signal separation is achieved.
It significantly improves the accuracy and efficiency of coal gangue sound signal separation, can accurately separate target sound in complex noise environments, has high audio separation accuracy, good robustness and generalization ability.
Smart Images

Figure CN120220715A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of coal gangue identification, and particularly relates to a device and method for separating coal gangue sound signals based on deep time-frequency feature fusion. Background Art
[0002] During the coal mining process, the identification and separation of coal and gangue are key links, and their accuracy and efficiency directly affect resource utilization rate and production safety. However, the coal mining process is filled with complex background noises, such as the noises generated when coal shearers, conveyors, and transfer machines are working, which pose great challenges to the separation of coal gangue sound signals.
[0003] In recent years, researchers have conducted extensive research on different coal gangue identification methods, such as natural gamma ray method, infrared detection method, image recognition method, and sound and vibration signal analysis method, etc. The above methods mostly have problems such as difficult installation and high cost. Due to the simplicity of sound signal acquisition and low implementation cost, relevant scholars have used it for coal gangue identification research. Although some researchers have used methods such as independent component analysis, wavelet packet transform, and mel cepstral coefficients for coal gangue identification, when dealing with complex coal mine background noises, traditional methods show limitations such as poor noise adaptability, signal distortion, and weak real-time processing ability. Summary of the Invention
[0004] The purpose of the present invention is to provide a device and method for separating coal gangue sound signals based on deep time-frequency feature fusion, which can effectively cope with complex background noises, significantly improve the accuracy and efficiency of coal gangue sound signal separation, and provide technical support for the intelligentization of coal mining.
[0005] To achieve the above purpose, the present invention provides a device for separating coal gangue sound signals based on deep time-frequency feature fusion, including a time-domain feature extraction module, a frequency-domain feature extraction module, and a feature fusion and source separation module;
[0006] The time-domain feature extraction module is configured to use a convolutional neural network, a bidirectional long short-term memory network, and a Transformer based on a multi-head attention mechanism to achieve the extraction of deep features and the capture of temporal information;
[0007] The frequency-domain feature extraction module is configured to convert the input time-domain signal into a frequency-domain representation;
[0008] The feature fusion and source separation module is configured to merge the features extracted from the time and frequency domains into a unified feature representation.
[0009] In addition, the present invention also mentions a method for separating coal gangue sound signals based on deep time-frequency feature fusion, which specifically includes the following steps;
[0010] Step 1: Build a dataset of coal gangue sound signals in a noise environment;
[0011] Simulate the single caving coal acoustic environment when each engineering equipment works alone, as well as the complex and mixed noise environment when various engineering equipments work simultaneously. Based on various individual sound events, build a dataset of coal gangue sound signals in a noise environment with different noise complexities;
[0012] Step 2: Extract time series features through the time domain feature extraction module;
[0013] Introduce a time masking mechanism, learn sample features from the missing data by randomly generating time masks, and use convolutional layers and bidirectional long short-term memory networks to deeply extract time domain features from the data. Subsequently, further extract time series features through the Transformer layer;
[0014] Step 3: Extract frequency domain sequence features through the frequency domain feature extraction module;
[0015] First, perform a fast Fourier transform on the original audio data to convert the samples to the frequency domain, then use a linear layer to process the magnitude of the FFT, and extract features in the frequency domain space through the Transformer layer;
[0016] Step 4: Introduce the Transformer encoder layer with a multi-head attention mechanism in the time domain and frequency domain feature modules, enabling the time domain feature extraction module and the frequency domain feature extraction module to simultaneously focus on multiple frequency subspaces;
[0017] Step 5: Reconstruct and separate signals through the feature fusion and source separation module;
[0018] Introduce a hybrid feature fusion mechanism to achieve deep fusion of the time domain and frequency domain. Subsequently, process the fused features through a decoding network composed of deconvolutional layers and mapping layers, and finally reconstruct and separate the signals;
[0019] Step 6: According to the caving coal sequence of the hydraulic support, perform manual caving coal operations on the remaining hydraulic supports in turn. Starting from the caving coal of the second hydraulic support, the collected sound signals are processed through Step 2 and Step 3 in turn, and then input into the trained DTFFNet (Deep Time-Frequency Feature Fusion) model to achieve fast and accurate separation of coal gangue sound signals.
[0020] Preferably, in Step 1, it specifically includes the following steps:
[0021] Step 1.1: For the noise generated when the on-site working equipment operates, use a high-sensitivity directional microphone array to collect the coal gangue sound signals and acoustic feature data generated by the operation of the equipment from multiple angles, and use sound insulation barriers and sound-absorbing materials to isolate and attenuate the surrounding environmental noise to restore a pure single caving coal acoustic scene;
[0022] The noise signals are respectively: the right cutting part of the shearer, the left cutting part of the shearer, the conveyor, the rear conveyor, and the front conveyor; the sound signals of coal and gangue are collected in the laboratory environment;
[0023] Step 1.2: Select multiple representative monitoring points in the coal mining operation area, and deploy omnidirectional audio acquisition devices; synchronously collect the sound signals generated by the joint operation of multiple devices, and record the operating parameters, working states, mutual relationships and collaborative operation modes of each device;
[0024] Step 1.3: By building a top coal caving simulation test bench and installing sound sensors, pure sound signals of coal and gangue are collected in the laboratory environment;
[0025] Step 1.4: Respectively use the noise signals during the operation of the right cutting part of the shearer, the left cutting part of the shearer, the conveyor, the rear conveyor, and the front conveyor as background noise and the sound signals of the coal and gangue falling states to be separated, and mix them one by one to form a noise environment coal and gangue separation data set, and convert it to mono;
[0026] Step 1.5: Preprocess the audio data by using the overlapping segmentation method;
[0027] Specifically as follows: By setting the length and step size of the sliding window, multiple overlapping short audio segments are generated from a single long audio; if an audio has S data points, the length of each segmented segment is M, the step size of the sliding window is L, and the number of generated samples N is calculated by the following formula:
[0028]
[0029] Each sample data X = {X (1) , X (2) , X (3) , …, X (n)} is represented as a series of consecutive segmented segments; the X(k) segment represents the k-th window data, which is represented as: X(k) = [x((k - 1)L), x((k - 1)L + 1), …, x((k - 1)L + M - 1)], and each X(k) is a data segment of length M extracted from the original data X with a step size of L.
[0030] Preferably, in step 2, it specifically includes the following steps:
[0031] Step 2.1: Input the sample data X into the convolutional layer, and map the one-dimensional audio signal to a high-dimensional feature space through a series of convolutional operations to extract the features of the local area; this operation is represented as:
[0032] Y = ReLU(W * X + b) (2);
[0033] Among them, W is the weight of the convolution kernel, b is the bias term, * represents the convolution operation, and ReLU is a non-linear activation function used to increase the expressive power of the model;
[0034] Step 2.2: After extracting the preliminary time features, use a bidirectional long short-term memory network to process the time series data in these feature maps; the bidirectional long short-term memory network can comprehensively capture the long-term dependence information in the audio signal by learning forward and backward information simultaneously; its basic mathematical model is:
[0035] (h t ,c t ) = BiLSTM(h t-1 ,c t-1 ,Y t ) (3);
[0036] Among them, h t and c t respectively represent the hidden state and cell state at time step t, and Y t is the input feature sequence;
[0037] Step 2.3: Adopt a Transformer layer based on the multi-head self-attention mechanism as another strategy for processing time features;
[0038] In the multi-head attention mechanism, each input element is linearly transformed to obtain the query Q, key K, and value V:
[0039] Q = XW Q ,K = XW K ,V = XW V (4);
[0040] Among them, X is the sample data, and W Q , W K , W V are the corresponding weight matrices;
[0041] Calculate the attention weights through the dot product of the query and the key:
[0042]
[0043] Among them, d k is the dimension of the key vector K;
[0044] The multi-head attention mechanism independently learns different feature subspace relationships through multiple heads, and splices the outputs of multiple attention heads to obtain the final output representation:
[0045] O = MultiHead(Q,K,V) = Concat(head1,…,headh )W o (6);
[0046] Through the multi - head attention mechanism, the model can process all position information in the sequence in parallel, and its operation is summarized as:
[0047] U = Transformer(H) (7);
[0048] Where, H represents the sequence output by the long - short - term memory network, and U is the output encoded by the Transformer layer.
[0049] Preferably, in step 2, the time masking process includes the following two steps:
[0050] Step S1: According to the preset masking strategy, determine the proportional value of the elements in the masking tensor that need to be set to 0. After determining the proportion, use a randomization algorithm to accurately and randomly distribute the corresponding number of 0 elements to the masking tensor with the same dimension as the original input sequence X, ensuring the randomness and uniformity of the masking positions, so that the model can comprehensively and evenly encounter the missing data situations at different positions during the training process. The element type of the masking tensor is a boolean binary value, and the positions of the 0 elements represent that the original data at the corresponding positions will be masked in this training;
[0051] Step S2: Perform an element - by - element dot - product operation between the generated masking tensor and the original time series X to obtain the masked sample sequence.
[0052] Preferably, in step 3, through the fast Fourier transform FFT, the input time - domain signal is converted into a frequency - domain representation. The FFT is based on the Fourier transform and decomposes a function or signal into a combination of sine waves and / or cosine waves of different frequencies; for discrete signals, the FFT is expressed as:
[0053]
[0054] Where, X[k] is the k - th frequency component in the frequency domain, x[n] is the n - th sample of the time - domain signal, N is the number of FFT points, and i is the imaginary unit;
[0055] This conversion transforms the time - domain sample x[n] into the frequency - domain sample X[k], and the frequency - domain sample X[k] contains the amplitude and phase information of frequency k;
[0056] After the FFT, the frequency - domain representation of the signal is input into a fully - connected layer to learn the complex relationships and patterns between frequency components; the mathematical representation of this layer is:
[0057] F out = W·X(k)+b (7);
[0058] Among them, F out is the output of the frequency domain layer, and W and b are the weight and bias of this layer respectively;
[0059] To further process the frequency domain information, the frequency domain features introduce the Transformer encoder layer of the multi-head attention mechanism. This layer can capture the complex dependencies between frequency components. The use of the multi-head attention mechanism allows the model to process the information flow between different frequencies in parallel, enabling the frequency domain feature extraction module to simultaneously focus on multiple frequency subspaces;
[0060] Each attention head learns independent frequency relationships, so as to be able to focus on different parts of the frequency signal. The specific expression is:
[0061] F = MultiHead(F out ) = Concat(head1,…,head h )W o (8);
[0062] Among them, MultiHead is the output of the multi-head attention, and W o is the final output transformation matrix, and Concat represents the concatenation operation.
[0063] Under the processing of the multi-head attention, the frequency domain feature extraction module can effectively separate different sound sources from the mixed signal, providing rich frequency feature information for subsequent feature fusion and sound source separation.
[0064] Preferably, in step 5, the feature fusion and source separation module utilizes the complementary information between different features to enhance the model's ability to identify and separate each source in the mixed signal; when fusing these features, the model concatenates the time and frequency domain features in the feature dimension and then integrates them together through a fusion layer;
[0065] Let the time feature matrix be T and the frequency domain feature matrix be F. The fusion operation is expressed as:
[0066] F = W f ·F + b f (9);
[0067] T = W t ·T + b t (10);
[0068] F fused = W c ·[T; F] + b c (11);
[0069] Among them, W f , W t, W c and b f , b t , b c are the weights and biases of their respective layers, and [T; F] represents concatenating the time and frequency domain features along a specific axis;
[0070] After the fused features are processed, these features are first passed through an output linear layer to predict the features of each source linear out denotes:
[0071] linear out = W o ·F fused + b o (12);
[0072] where, W o is the weight of the linear layer, and b o is the bias of the linear layer;
[0073] Then these features are converted back to the time domain through a decoder to generate the final audio signal for each source; the decoder is a deconvolution network that reconstructs each individual source based on the predicted mask or source features:
[0074] S i = Decoder(D i ) (13);
[0075] where, D i is the feature of the i-th source output from the deconvolution layer, and S i is the reconstructed i-th source signal.
[0076] The beneficial technical effects brought by the present invention:
[0077] By fusing the time domain and frequency domain features, the present invention effectively improves the processing ability of complex audio signals, and can accurately separate the target sound in the complex coal gangue audio signal environment; DTFFNet has high audio separation accuracy under various signal-to-noise ratio conditions, especially in the high signal-to-noise ratio environment, the separation performance of the model is significantly improved; compared with the traditional method that only relies on time domain feature extraction, DTFFNet shows more significant performance advantages when dealing with complex and multi-source noise signals, reflecting its good robustness and generalization ability; in addition, the verification of DTFFNet model in multiple scenarios shows that DTFFNet is not only applicable to the sound signal separation task in the coal mine operation environment, but also can show its application potential in a wide range of audio processing scenarios; the results of this study provide an efficient and accurate coal gangue sound signal separation technology for the coal mine industry, which is expected to promote the further development of intelligent mine technology, and at the same time provides new research ideas and technical means for the field of multi-source audio signal processing. Description of the Drawings
[0078] Figure 1 is a flowchart of the method for separating coal gangue sound signals according to the present invention;
[0079] Figure 2 is a schematic diagram of the design of the sound signal acquisition platform according to the present invention;
[0080] Figure 3 is the structure diagram of the DTFFNet model in the present invention;
[0081] Figure 4 is the structure diagram of feature fusion and source separation in the present invention. Detailed Embodiment
[0082] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0083] The present invention aims to provide a method (DTFFNet) for separating coal gangue sound signals based on deep time-frequency feature fusion, which can effectively cope with complex background noise, significantly improve the accuracy and efficiency of coal gangue sound signal separation, and provide technical support for the intelligentization of coal mining.
[0084] To achieve the above object, the present invention provides a method for separating coal gangue sound signals based on deep time-frequency feature fusion, the process of which is as Figure 1 shown, including the following steps;
[0085] Step 1: Simulate a single caving coal acoustic environment when each engineering equipment works alone, as well as a complex and mixed noise environment when various engineering equipments work simultaneously. On the basis of various individual sound events, build a coal gangue sound signal dataset with different noise complexities;
[0086] Step 2: The time-domain feature extraction module learns sample features from the missing data by randomly generating time masks, and uses convolutional layers and bidirectional LSTM networks to deeply extract time-domain features of the data. Subsequently, time series features are further extracted through the Transformer layer;
[0087] Step 3: The frequency-domain feature extraction module first performs a fast Fourier transform (FFT) on the original audio data to transform the samples into the frequency domain, then uses a linear layer to process the amplitude of the FFT, and extracts features in the frequency domain space through the Transformer layer;
[0088] Step 4: In order to further process time-domain and frequency-domain information, a Transformer encoder layer with a multi-head attention mechanism is introduced into the time-domain and frequency-domain feature modules, so that the time-domain and frequency-domain feature extraction modules can simultaneously focus on multiple frequency subspaces;
[0089] Step 5: A hybrid feature fusion mechanism is introduced to achieve deep fusion of time domain and frequency domain. The fused features are then processed through a decoding network composed of deconvolution layers and mapping layers to finally reconstruct the separated signals.
[0090] Step 6: According to the order of hydraulic support coal placement, the remaining hydraulic supports are manually placed in turn, starting from the second hydraulic support. The collected sound signals are processed in steps 2 and 3 and then input into the trained DTFFNet model (such as Figure 3 As shown in the figure, the rapid and accurate separation of coal gangue acoustic signals is achieved.
[0091] Furthermore, the coal gangue acoustic signal separation method of the sound signal introduces a time masking mechanism, which aims to effectively reduce the model's excessive dependence on specific data points by forcing the model to learn sample features from missing data, thereby enhancing its ability to capture and learn key features.
[0092] Specifically, in the training process of the DTFFNet model, the temporal masking process mainly covers the following two key steps:
[0093] First, according to the preset masking strategy, determine the proportion of elements in the mask tensor that need to be set to 0. The proportion can be flexibly adjusted and optimized according to specific training requirements, data characteristics, and model performance optimization goals. After determining the proportion, the corresponding number of 0 elements is accurately and randomly assigned to the mask tensor with exactly the same dimensions as the original input sequence X through a randomization algorithm to ensure the randomness and uniformity of the mask position, so that the model can be fully and evenly exposed to missing data at different positions during training. The element type of the mask tensor is a Boolean binary value, and the position of the 0 element represents that the original data at the corresponding position will be masked in this training;
[0094] Secondly, perform element-by-element point multiplication on the generated mask tensor and the original time series X to obtain a sample sequence after masking. Through this operation, the input data received by the model during the training process will show a state of partial data missing, which will prompt the model to actively mine the potential key features in the data to make up for the information loss that may be caused by data missing, so that the model has stronger robustness and adaptability when facing incomplete data that may appear in practical applications, especially in the audio source separation task, it can focus on significant features more accurately and improve the separation effect and accuracy;
[0095] To simplify the subsequent description and the description of the model operation process, in the subsequent content of the present invention, the symbol X is uniformly used to refer to the model input sample after masking processing, so as to clearly and concisely elaborate on key technical details such as the overall architecture of the model, the training process, and the performance optimization method, and ensure that the present invention has good readability, operability, and scalability in terms of technical implementation and application promotion.
[0096] Furthermore, for the method of separating coal gangue sound signals from sound signals, in order to be able to simulate the single caving coal acoustic environment when each engineering device works alone, as well as the complex and mixed noise environment when various engineering devices work simultaneously, a multi-source audio acquisition system is adopted to comprehensively and accurately collect sound data under different working conditions, and a data set of coal gangue sound signals in a noise environment is constructed. The specific steps are as follows:
[0097] S11: For the noise generated when the on-site operating equipment works, a high-sensitivity directional microphone array is used to collect the coal gangue sound signals and acoustic feature data generated by the operation of the equipment from multiple angles, and sound insulation barriers and sound-absorbing materials are used to isolate and attenuate the surrounding environmental noise to restore a pure single caving coal acoustic scene;
[0098] The noise signals are respectively: the right cutting part of the shearer, the left cutting part of the shearer, the conveyor, the rear conveyor, and the front conveyor. Then, the sound signals of coal and gangue are collected in the laboratory environment.
[0099] S12: Select multiple representative monitoring points in the coal mining operation area and deploy omnidirectional audio acquisition equipment; synchronously collect the sound signals generated by the joint operation of multiple devices, and record the operating parameters, working states, mutual relationships, and collaborative operation modes of each device, providing a basis for subsequent analysis of the contribution ratio of different device sound sources to the mixed noise environment and the acoustic feature superposition effect.
[0100] S13: By building a caving coal simulation test bench and installing sound sensors, as Figure 2 shown, pure sound signals of coal and gangue are collected in the laboratory environment;
[0101] S14: The noise signals when the right cutting part of the shearer, the left cutting part of the shearer, the conveyor, the rear conveyor, and the front conveyor are working are respectively used as background noises and are mixed one by one with the sound signals of the coal and gangue falling states to be separated to form a noise environment coal gangue separation data set, and it is converted into a mono channel;
[0102] S15: In order to enhance data diversity and increase the number of samples, an overlapping segmentation method is adopted to preprocess the audio data;
[0103] Specifically, by setting the length and step size of the sliding window, multiple overlapping short audio segments can be generated from a single long audio; if an audio has S data points, the length of each segmented segment is M, and the step size of the sliding window is L, then the number of samples N that can be generated can be calculated by the following formula:
[0104]
[0105] Under this setting, each sample X = {X (1) , X (2) , X (3) , …, X (n)} can be represented as a series of consecutive segmented segments. Among them, the X(k) segment represents the k-th window data and can be expressed as: X(k) = [x((k - 1)L), x((k - 1)L + 1), …, x((k - 1)L + M - 1)], and each X(k) is a data segment of length M extracted from the original data X with a step size of L; this method not only increases the number of samples but also simulates more diverse auditory scenarios by introducing overlapping segments, which helps to improve the adaptability and generalization ability of the model to different audio environments.
[0106] Furthermore, the DTFFNet model of the coal gangue identification sensor and identification method for sound signals mainly consists of a time-domain feature extraction module, a frequency-domain feature extraction module, and a feature fusion and source separation module.
[0107] Specifically, the time-domain feature extraction module mainly uses a convolutional neural network (CNN), a bidirectional long short-term memory network (BiLSTM), and a Transformer based on a multi-head attention mechanism to achieve the extraction of deep features and the capture of temporal information. The sample data X is first input to the convolutional layer, and through a series of convolutional operations, the one-dimensional audio signal is mapped to a high-dimensional feature space to extract the features of the local region. This operation can be expressed as:
[0108] Y = ReLU(W * X + b) (2);
[0109] where W is the weight of the convolutional kernel, b is the bias term, * represents the convolutional operation, and ReLU is a non-linear activation function used to increase the expressive power of the model;
[0110] After extracting the preliminary time features, bidirectional LSTM is used to process the time series data in these feature maps. Bidirectional LSTM can comprehensively capture the long-term dependence information in the audio signal by learning forward and backward information simultaneously. Its basic mathematical model is:
[0111] (h t , c t ) = BiLSTM(ht-1 , c t-1 , Y t ) (3);
[0112] Here, h t and c t represent the hidden state and cell state at time step t respectively, and Y t is the input feature sequence.
[0113] Although bidirectional LSTM can effectively handle the dependency problem of time series, its computational complexity is high, and there are still certain challenges in processing very long sequences. Therefore, the model adopts the Transformer layer based on the multi-head self-attention mechanism as another strategy for time feature processing. The multi-head attention mechanism allows the model to process the information at all positions in the sequence in parallel and learn the information of different parts through multiple heads, enabling the model to capture long-range dependencies more flexibly.
[0114] In the multi-head attention mechanism, each input element is linearly transformed to obtain query (Query), key (Key), and value (Value):
[0115] Q = XW Q , K = XW K , V = XW V (4);
[0116] where X is the input feature, and W Q , W K , W V are the corresponding weight matrices. Then the dot product of the query and the key is used to calculate the attention weights:
[0117]
[0118] Here, d k is the dimension of the key vector K. The multi-head attention mechanism independently learns different feature subspace relationships through multiple heads, and further concatenates the outputs of multiple attention heads to obtain the final output representation:
[0119] O = MultiHead(Q, K, V) = Concat(head1,..., head h )W o (6);
[0120] Through the multi-head attention mechanism, the model can process all position information in the sequence in parallel, significantly improving the ability to handle long-range dependencies and enhancing the efficiency of feature capture. Its operation can be summarized as:
[0121] U = Transformer(H) (7);
[0122] Among them, H represents the sequence output by the LSTM, and U is the output encoded by the Transformer layer.
[0123] Specifically, for the frequency-domain feature extraction module, the first step is to convert the input time-domain signal into a frequency-domain representation; this is achieved through the Fast Fourier Transform (FFT). The FFT is based on the Fourier transform, which decomposes a function or signal into a combination of sine waves and / or cosine waves of different frequencies. For a discrete signal, the FFT can be expressed as:
[0124]
[0125] where X[k] is the k-th frequency component in the frequency domain, x[n] is the n-th sample of the time-domain signal, N is the number of FFT points, usually taken as a power of 2 to optimize the calculation efficiency, and i is the imaginary unit.
[0126] This conversion transforms the time-domain samples x[n] into frequency-domain samples X[k], which contain the amplitude and phase information of frequency k.
[0127] After the FFT, the frequency-domain representation of the signal is input into a fully connected layer (linear layer) to learn the complex relationships and patterns between frequency components. The mathematical representation of this layer is:
[0128] F out = W·X(f) + b (7);
[0129] Here, F out is the output of the frequency-domain layer, and W and b are the weights and biases of this layer, respectively.
[0130] To further process the frequency-domain information, the frequency-domain features usually need to introduce the Transformer encoder layer with the multi-head attention mechanism. This layer can capture the complex dependencies between frequency components. The use of the multi-head attention mechanism allows the model to process the information flow between different frequencies in parallel, enabling the frequency-domain feature extraction module to simultaneously focus on multiple frequency subspaces;
[0131] Each attention head learns independent frequency relationships, enabling it to focus on different parts of the frequency signal. The specific expression is:
[0132] F = MultiHead(F out ) = Concat(head1,…,head h )W o (8);
[0133] Under the processing of multi-head attention, the frequency-domain feature extraction module can effectively separate different sound sources from the mixed signal, providing rich frequency feature information for subsequent feature fusion and sound source separation.
[0134] Specifically, the feature fusion and source separation module combines the features extracted from different processing paths (time and frequency domain) into a unified feature representation to more effectively reconstruct independent audio sources. The purpose of this fusion strategy is to utilize the complementary information between different features and enhance the model's ability to identify and separate each source in the mixed signal. When fusing these features, the model concatenates the time and frequency domain features along the feature dimension and then integrates them together through a fusion layer (linear layer). The main role of this fusion layer is to learn how to most effectively combine these features, thereby improving the source separation effect. As Figure 4 shown.
[0135] Let the time feature matrix be T and the frequency domain feature matrix be F. The fusion operation can be expressed as:
[0136] F = W f ·F + b f (9);
[0137] T = W t ·T + b t (10);
[0138] F fused = W c ·[T; F] + b c (11);
[0139] Among them, W f , W t , W c and b f , b t , b c are the weights and biases of their respective layers, and [T; F] represents concatenating the time and frequency domain features along a specific axis.
[0140] After the fused features are processed, these features first pass through an output linear layer to predict the feature representation of each source:
[0141] linear out = W o ·F fused + b o (12);
[0142] Then, a decoder is used to convert these features back to the time domain to generate the final audio signal for each source. The decoder is a deconvolution network that reconstructs each independent source based on the predicted mask or source features:
[0143] Si = Decoder(D i ) (13);
[0144] where D i is the feature of the i-th source output from the transposed convolutional layer, and S i is the reconstructed i-th source signal.
[0145] DTFFNet improves the robustness of the model through the time masking mechanism, and the deep fusion of joint time-frequency features significantly enhances the source separation ability of the model. The design of this architecture effectively improves the robustness and interpretability of the model while ensuring the accuracy of feature capture.
[0146] As an optimization, in step one, a multi-source audio acquisition system is proposed to comprehensively and accurately collect sound data under different working conditions, and construct a dataset of coal gangue sound signals in a noisy environment;
[0147] As an optimization, in step five, a feature fusion and source separation strategy is proposed to utilize the complementary information between different features to enhance the model's ability to identify and separate each source in the mixed signal;
[0148] Certainly, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by those skilled in the art within the scope of the essence of the present invention should also fall within the protection scope of the present invention.
Claims
1. A coal gangue sound signal separation device based on deep time-frequency feature fusion, characterized in that: It includes time domain feature extraction module, frequency domain feature extraction module, feature fusion and source separation module; The time domain feature extraction module is configured to use convolutional neural networks, bidirectional long short-term memory networks, and Transformers based on multi-head attention mechanisms to extract deep features and capture temporal information; A frequency domain feature extraction module is configured to convert an input time domain signal into a frequency domain representation; The feature fusion and source separation module is configured to merge the features extracted from the time and frequency domains into a unified feature representation.
2. A method for separating coal gangue acoustic signals based on deep time-frequency feature fusion, characterized in that: The coal gangue sound signal separation device based on deep time-frequency feature fusion as claimed in claim 1 specifically comprises the following steps: Step 1: Build a data set of coal gangue sound signals in a noisy environment; Simulate the single top coal caving acoustic environment when each engineering equipment works alone, and the complex, mixed noise environment when various engineering equipment works simultaneously. On the basis of various individual sound events, build a coal gangue sound signal dataset with different noise complexity in the noise environment; Step 2: Extract time series features through the time domain feature extraction module; A time mask mechanism is introduced to learn sample features from missing data by randomly generating time masks, and convolutional layers and bidirectional long short-term memory networks are used to perform in-depth time domain feature extraction on the data, followed by further extracting time series features through the Transformer layer. Step 3: Extract frequency domain sequence features through the frequency domain feature extraction module; First, the original audio data is subjected to a fast Fourier transform to convert the samples into the frequency domain. Then, the amplitude of the FFT is processed using a linear layer, and feature extraction is performed in the frequency domain through a Transformer layer. Step 4: Introduce the Transformer encoder layer with multi-head attention mechanism in the time domain and frequency domain feature modules, so that the time domain feature extraction module and the frequency domain feature extraction module can focus on multiple frequency subspaces at the same time; Step 5: Reconstruct the separated signal through feature fusion and source separation module; A hybrid feature fusion mechanism is introduced to achieve deep fusion of time domain and frequency domain. The fused features are then processed through a decoding network consisting of a deconvolution layer and a mapping layer to finally reconstruct the separation signal. Step 6: According to the order of coal placing of hydraulic supports, manual coal placing operation is performed on the remaining hydraulic supports in turn. Starting from the coal placing of the second hydraulic support, the collected sound signals are processed in steps 2 and 3 in turn, and then input into the trained deep time-frequency feature fusion model to achieve fast and accurate separation of coal gangue sound signals.
3. The method for separating coal gangue sound signals based on deep time-frequency feature fusion according to claim 2 is characterized in that: Step 1 specifically includes the following steps: Step 1.1: In view of the noise generated by the working equipment on site, a high-sensitivity directional microphone array is used to collect the coal gangue sound signals and acoustic characteristic data generated by the operation of the equipment from multiple angles, and the surrounding environmental noise is isolated and attenuated using sound insulation barriers and sound-absorbing materials to restore the pure single top coal caving acoustic scene; The noise signals are: the right cutting part of the coal mining machine, the left cutting part of the coal mining machine, the transfer machine, the rear conveyor, and the front conveyor; The acoustic signals of coal and gangue were collected in a laboratory environment; Step 1.2: Select multiple representative monitoring points in the coal mining operation area and deploy all-round audio collection equipment; Synchronously collect the sound signals generated by the joint operation of multiple devices, and record the operating parameters, working status, mutual relationship and collaborative operation mode of each device; Step 1.3: By building a top coal caving simulation test bench and installing sound sensors, the sound signals of pure coal and gangue are collected in a laboratory environment; Step 1.4: The noise signals of the right cutting part, the left cutting part, the transfer machine, the rear conveyor and the front conveyor when the coal mining machine is working are respectively used as background noise and mixed with the sound signals of the falling coal and gangue states to be separated one by one to form a noise environment coal-gangue separation data set, and converted into a mono channel; Step 1.5: Preprocess the audio data using overlapping segmentation method; The details are as follows: by setting the length and step size of the sliding window, multiple overlapping short audio segments are generated from a single long audio; if an audio has S data points, the length of each segment is M, the step size of the sliding window is L, and the number of samples N generated is calculated by the following formula: Each sample data X={X (1) ,X (2) ,X (3) ,…,X (n) } is represented as a series of continuous segmented segments; the X(k) segment represents the kth window data, expressed as: X(k) = [x((k-1)L), x((k-1)L+1), …, x((k-1)L+M-1)], where each X(k) is a data segment of length M extracted from the original data X with a step size of L.
4. The method for separating coal gangue acoustic signals based on deep time-frequency feature fusion according to claim 3 is characterized in that: Step 2 specifically includes the following steps: Step 2.1: Input the sample data X into the convolution layer, and map the one-dimensional audio signal to a high-dimensional feature space through a series of convolution operations to extract the features of the local area; this operation is expressed as: Y = ReLU(W*X+b) (2); Among them, W is the weight of the convolution kernel, b is the bias term, * represents the convolution operation, and ReLU is a nonlinear activation function used to increase the expressiveness of the model; Step 2.2: After extracting the preliminary time features, use the bidirectional long short-term memory network to process the time series data in these feature maps; the bidirectional long short-term memory network can fully capture the long-term dependency information in the audio signal by learning both forward and reverse information at the same time; its basic mathematical model is: (h t ,c t )=BiLSTM(h t-1 ,c t-1 ,Y t ) (3); Among them, h t and c t denote the hidden state and cell state at time step t, respectively, and Y t is the input feature sequence; Step 2.3: Use the Transformer layer based on the multi-head self-attention mechanism as another strategy for temporal feature processing; In the multi-head attention mechanism, each input element is linearly transformed to obtain the query Q, key K and value V: Q=XW Q ,K=XW K ,V=XW V (4); Among them, X is the sample data, W Q , W K , W V is the corresponding weight matrix; The attention weight is calculated by taking the dot product of the query and the key: Among them, d k is the dimension of the key vector K; The multi-head attention mechanism uses multiple heads to independently learn different feature subspace relationships, and concatenates the outputs of multiple attention heads to obtain the final output representation: O=MultiHead(Q,K,V)=Concat(head1,…,head h )W o (6); Through the multi-head attention mechanism, the model can process all position information in the sequence in parallel. Its operation can be summarized as follows: U = Transformer(H) (7); Among them, H represents the sequence output by the long short-term memory network, and U is the output after Transformer layer encoding.
5. The method for separating coal gangue sound signals based on deep time-frequency feature fusion according to claim 2, characterized in that: In step 2, the time masking process includes the following two steps: Step S1: According to the preset masking strategy, determine the proportion of elements in the mask tensor that need to be set to 0. After determining the proportion, use a randomization algorithm to accurately and randomly assign the corresponding number of 0 elements to the mask tensor with exactly the same dimension as the original input sequence X, ensuring the randomness and uniformity of the mask position, so that the model can be fully and evenly exposed to missing data at different positions during training. The element type of the mask tensor is a Boolean binary value, and the position of the 0 element represents that the original data at the corresponding position will be masked in this training; Step S2: Perform an element-by-element dot multiplication operation on the generated mask tensor and the original time series X to obtain a sample sequence after mask processing.
6. The method for separating coal gangue sound signals based on deep time-frequency feature fusion according to claim 2, characterized in that: In step 3, the input time domain signal is converted into frequency domain representation by fast Fourier transform FFT. FFT is based on Fourier transform, which decomposes a function or signal into a combination of sine waves and / or cosine waves of different frequencies. For discrete signals, FFT is expressed as: Where X[k] is the kth frequency component in the frequency domain, x[n] is the nth sample of the time domain signal, N is the number of FFT points, and i is the imaginary unit; This transformation converts the time domain sample x[n] into the frequency domain sample X[k], which contains the amplitude and phase information of frequency k; After the FFT, the frequency domain representation of the signal is input into a fully connected layer for learning the complex relationships and patterns between the frequency components; the mathematical representation of this layer is: F out =W·X(k)+b (7); Among them, F out is the output of the frequency domain layer, W and b are the weight and bias of this layer respectively; To further process the frequency domain information, the frequency domain features introduce a Transformer encoder layer with a multi-head attention mechanism, which can capture the complex dependencies between frequency components. The use of the multi-head attention mechanism allows the model to process information flows between different frequencies in parallel, so that the frequency domain feature extraction module can focus on multiple frequency subspaces at the same time. Each attention head learns an independent frequency relationship, so that it can focus on different parts of the frequency signal. The specific expression is: F=MultiHead(F out )=Concat(head1,…,head h )W o (8); Among them, MultiHead is the output of multi-head attention, W o is the final output transformation matrix, Concat represents the concatenation operation; Under the processing of multi-head attention, the frequency domain feature extraction module can effectively separate different sound sources from the mixed signal, providing rich frequency feature information for subsequent feature fusion and sound source separation.
7. The method for separating coal gangue sound signals based on deep time-frequency feature fusion according to claim 2, characterized in that: In step 5, the feature fusion and source separation module uses the complementary information between different features to enhance the model's ability to identify and separate each source in the mixed signal; when fusing these features, the model concatenates the time and frequency domain features in the feature dimension and then integrates them together through a fusion layer; Assume that the time feature matrix is T and the frequency domain feature matrix is F, and the fusion operation is expressed as: F=W f ·F+b f (9); T=W t ·T+b t (10); F fused =W c ·[T;F]+b c (11); Among them, W f , W t , W c and b f , b t , b c are the weights and biases of their respective layers, respectively. [T; F] represents the concatenation of time and frequency domain features along a specific axis; After the fused features are processed, they are first passed through an output linear layer to predict the features of each source. out express: linear out =W o ·F fused +b o (12); Among them, W o is the weight of the linear layer, b o The paranoia of the linear layer; These features are then transformed back into the time domain to produce the final audio signal for each source via a decoder, a deconvolutional network that reconstructs each individual source based on the predicted mask or source features: S i =Decoder(D i ) (13); Among them, D i is the feature of the i-th source output from the deconvolution layer, S i is the reconstructed i-th source signal.