Transformer substation noise intelligent separation method, device, equipment and medium
By using a multi-channel voice enhancement model and a complex adaptive beamforming network to separate substation noise, the problem of noise aliasing within the substation was solved, thereby improving communication quality and fault early warning capabilities.
Patent Information
- Application Number
- CN202511863761.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-20
AI Technical Summary
The high-intensity and complex acoustic environment inside substations causes background noise to be severely mixed with the voices of maintenance personnel and abnormal equipment sounds. Existing technologies are unable to effectively separate the noise, affecting communication quality and fault early warning capabilities.
A multi-channel speech enhancement model is adopted. By collecting and preprocessing device noise data, a mixed noise dataset is constructed, the multi-channel speech enhancement model is optimized, and noise separation is performed using a complex adaptive beamforming network and a complex fully convolutional network.
It improves noise separation performance, enhances communication quality, and increases the accuracy of fault warning.
Smart Images

Figure CN121708950A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of audio processing, in particular to a substation noise intelligent separation method, device, equipment and medium. BACKGROUND
[0002] As the core hub of the power system, the operation and maintenance of the substation highly depend on the voice communication of the on-site personnel, the acoustic monitoring of the equipment state, and the intelligent inspection based on voice instructions. However, the interior of the substation is a typical high-intensity and high-complexity acoustic environment, the background noise and the voice of the operation and maintenance personnel and the abnormal sound of the equipment are seriously overlapped, which causes the key audio information to be submerged, directly affecting the communication quality, the accuracy of the automatic instruction recognition, and the early fault warning ability based on sound.
[0003] In the prior art, noise separation is usually performed by a digital signal processing filtering method and an acoustic feature separation method based on traditional machine learning, but the former has weak effect on transient features during processing, and the latter is difficult to separate similar spectral components, and neither can achieve ideal noise separation effect. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a substation noise intelligent separation method, device, equipment and medium, which can process noise signals through a multi-channel voice enhancement model, thereby improving the effect of noise separation. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses a substation noise intelligent separation method, comprising:
[0006] Collecting device noise generated by a plurality of target devices working in a target environment, and preprocessing the device noise to construct a target mixed noise data set based on the obtained preprocessed noise data; the target mixed noise data set is a data set simulating substation noise;
[0007] Mixing the original audio signal to obtain a mixed audio signal, decomposing the mixed audio signal through a preset sound separation model to obtain a decomposed audio signal, and then constructing a signal distortion ratio and a scale invariant signal distortion ratio through the original audio signal, the decomposed audio signal and the mixed audio signal;
[0008] Optimizing a preset multi-channel voice enhancement model through the signal distortion ratio, the scale invariant signal distortion ratio and the target mixed noise data set to obtain a target multi-channel voice enhancement model;
[0009] Performing noise separation on the noise data to be processed based on the target multi-channel voice enhancement model to obtain a separated noise signal and a separated audio signal.
[0010] Optionally, the device noise generated by the target devices in the target environment is collected, and the device noise is preprocessed to construct a target mixed noise dataset based on the obtained preprocessed noise data, comprising:
[0011] In the target environment, the device working states corresponding to the fault states and the load conditions of the target devices are simulated, and the device noise generated by the target devices in the device working states is collected by the preset microphone array; the device noise includes noise generated by a single target device and noise generated by no less than one target device;
[0012] The noise data with a time length less than a preset time length threshold in the device noise is removed to obtain removed noise data, and the removed noise data and a preset background noise signal are subjected to Hilbert transform to obtain a complex time domain analytic signal; the preset background noise signal is a pre-collected substation background noise;
[0013] The target mixed noise dataset is constructed based on the complex time domain analytic signal.
[0014] Optionally, a signal distortion ratio is constructed based on the original audio signal, the decomposed audio signal and the mixed audio signal, comprising:
[0015] A difference between the original audio signal and the mixed audio signal is determined to obtain a first difference, and a difference between the original audio signal and the decomposed audio signal is determined to obtain a second difference;
[0016] A ratio between a square of a norm of the mixed audio signal and a square of a norm of the first difference is determined to obtain a first ratio, and a ratio between a square of a norm of the original audio signal and a square of a norm of the second difference is determined to obtain a second ratio;
[0017] Ten times of base-ten logarithmic operation is performed on the first ratio to obtain a first operation result, and ten times of base-ten logarithmic operation is performed on the second ratio to obtain a second operation result;
[0018] A difference between the first operation result and the second operation result is taken as a signal distortion ratio.
[0019] Optionally, a scale-invariant signal distortion ratio is constructed based on the original audio signal, the decomposed audio signal and the mixed audio signal, comprising:
[0020] An inner product of the mixed audio signal and the original audio signal is determined to obtain a first inner product, and a ratio between the first inner product and a square of a norm of the original audio signal is determined to obtain a first scalar.
[0021] multiplying the second scalar and the original audio signal to obtain a second processed audio signal, and calculating a difference between the decomposed audio signal and the second processed audio signal to obtain a second signal difference;
[0022] calculating a ratio between a square of a norm of the second processed audio signal and a square of a norm of the second signal difference to obtain a fourth ratio;
[0023] determining an inner product between the decomposed audio signal and the original audio signal to obtain a second inner product, and determining a ratio between the second inner product and a square of a norm of the original audio signal to obtain a second scalar;
[0024] multiplying the second scalar and the original audio signal to obtain a second processed audio signal, and calculating a difference between the decomposed audio signal and the second processed audio signal to obtain a second signal difference;
[0025] calculating a ratio between a square of a norm of the second processed audio signal and a square of a norm of the second signal difference to obtain a fourth ratio;
[0026] performing a ten-base logarithm operation on the third ratio by ten times to obtain a third operation result, and performing a ten-base logarithm operation on the fourth ratio by ten times to obtain a fourth operation result;
[0027] taking a difference between the third operation result and the fourth operation result as a scale-invariant signal distortion ratio.
[0028] Optionally, the preset multi-channel speech enhancement model is optimized by the signal distortion ratio, the scale-invariant signal distortion ratio, and the target mixed noise data set to obtain a target multi-channel speech enhancement model, which comprises:
[0029] a loss function of the preset multi-channel speech enhancement model is constructed based on the signal distortion ratio;
[0030] the preset multi-channel speech enhancement model is trained by the target mixed noise data set to obtain a to-be-confirmed multi-channel speech enhancement model;
[0031] the to-be-confirmed multi-channel speech enhancement model is verified in performance based on the signal distortion ratio and the scale-invariant signal distortion ratio;
[0032] if a model performance of the to-be-confirmed multi-channel speech enhancement model reaches a preset performance condition, the to-be-confirmed multi-channel speech enhancement model is taken as a target multi-channel speech enhancement model;
[0033] If the model performance of the to-be-confirmed multi-channel speech enhancement model does not reach the preset performance condition, jump to the step of training the preset multi-channel speech enhancement model through the target mixed noise data set to obtain a to-be-confirmed multi-channel speech enhancement model until the model performance of the to-be-confirmed multi-channel speech enhancement model reaches the preset performance condition.
[0034] Optionally, the preset multi-channel speech enhancement model is a model composed of a complex adaptive beamforming network and a complex fully convolutional network.
[0035] Optionally, the noise separation based on the target multi-channel speech enhancement model on the to-be-processed noise data to obtain the separated noise signal and the separated audio signal comprises:
[0036] convolving the to-be-processed noise through the complex adaptive beamforming network in the target multi-channel speech enhancement model to obtain an enhanced noise signal;
[0037] calculating a time-domain complex ratio mask of the enhanced noise signal through the complex fully convolutional network, and separating the to-be-processed noise through the time-domain complex ratio mask to obtain the separated noise signal and the separated audio signal.
[0038] In a second aspect, the present application discloses a substation noise intelligent separation device, comprising:
[0039] A data set construction module is configured to collect device noise generated by a plurality of target devices in a target environment during operation, and pre-process the device noise to construct a target mixed noise data set based on obtained pre-processed noise data; the target mixed noise data set is a data set simulating substation noise.
[0040] A distortion parameter calculation module is configured to mix an original audio signal to obtain a mixed audio signal, decompose the mixed audio signal through a preset sound separation model to obtain a decomposed audio signal, and then construct a signal distortion ratio and a scale-invariant signal distortion ratio through the original audio signal, the decomposed audio signal, and the mixed audio signal.
[0041] A model optimization module is configured to optimize a preset multi-channel speech enhancement model through the signal distortion ratio, the scale-invariant signal distortion ratio, and the target mixed noise data set to obtain a target multi-channel speech enhancement model.
[0042] A noise separation module is configured to separate noise from to-be-processed noise data based on the target multi-channel speech enhancement model to obtain a separated noise signal and a separated audio signal.
[0043] In a third aspect, the present application discloses an electronic device, comprising:
[0044] a memory for storing a computer program;
[0045] a processor for executing the computer program to implement the intelligent substation noise separation method as described above.
[0046] In a fourth aspect, the present application discloses a computer readable storage medium for storing a computer program, wherein the computer program is executed by a processor to implement the intelligent substation noise separation method as described above.
[0047] In the present application, device noise generated by a plurality of target devices in a target environment can be collected, and the device noise can be preprocessed to construct a target mixed noise data set based on the obtained preprocessed noise data; the target mixed noise data set is a data set simulating substation noise; an original audio signal is mixed to obtain a mixed audio signal, and the mixed audio signal is decomposed by a pre-set sound separation model to obtain a decomposed audio signal, and then a signal distortion ratio and a scale invariant signal distortion ratio are constructed by the original audio signal, the decomposed audio signal and the mixed audio signal; a pre-set multi-channel speech enhancement model is optimized by the signal distortion ratio, the scale invariant signal distortion ratio and the target mixed noise data set to obtain a target multi-channel speech enhancement model; and noise separation is performed on the target mixed noise data set based on the target multi-channel speech enhancement model to obtain a separated noise signal and a separated audio signal. As can be seen, the device noise generated by the target devices in the target environment can be collected by the method of the present application, and then the device noise is preprocessed to construct a target mixed noise data set based on the obtained preprocessed noise data, and then the original audio signal needs to be mixed, and the mixed audio signal needs to be decomposed to obtain a decomposed audio signal, and then a signal distortion ratio and a scale invariant signal distortion ratio are constructed by the original audio signal, the decomposed audio signal and the mixed audio signal. And the pre-set multi-channel speech enhancement model needs to be optimized according to the above distortion ratio and the target mixed noise data set, so as to separate the noise based on the obtained target multi-channel speech enhancement model, to obtain a separated noise signal and a separated audio signal. In this way, the noise signal can be processed by the improved multi-channel speech enhancement model, and the effect of noise separation is improved. BRIEF DESCRIPTION OF DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description only constitute a part of the embodiments of the present application, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of the provided drawings.
[0049] Figure 1 A substation noise intelligent separation method flow chart disclosed by the present application;
[0050] Figure 2 A CNABTCN network structure schematic diagram disclosed by the present application;
[0051] Figure 3 A complex LSTM rule schematic diagram disclosed by the present application;
[0052] Figure 4 A local complex TCN network structure schematic diagram disclosed by the present application;
[0053] Figure 5 A substation noise intelligent separation device structure schematic diagram disclosed by the present application;
[0054] Figure 6 A structure diagram of an electronic device disclosed by the present application. DETAILED DESCRIPTION
[0055] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without any creative effort belong to the protection scope of the present application.
[0056] In the prior art, noise separation is usually performed through a digital signal processing filtering method and an acoustic feature separation method based on a traditional machine learning, but the former has weak effect on transient characteristics during processing, and the latter is difficult to separate similar spectral components, and neither can achieve ideal noise separation effect.
[0057] In order to overcome the above technical problems, the present application discloses a substation noise intelligent separation method, device, equipment and medium, which can process noise signals through a multi-channel speech enhancement model, thereby improving the effect of noise separation.
[0058] Referring to Figure 1 The embodiments of the present application disclose a substation noise intelligent separation method, which comprises:
[0059] In step S11, the device noise generated by the target devices in the target environment is collected, and the device noise is preprocessed to construct a target mixed noise dataset based on the obtained preprocessed noise data; the target mixed noise dataset is a dataset simulating substation noise.
[0060] In this embodiment, the device noise generated by the target devices in the target environment needs to be collected, and then the device noise is preprocessed, and finally a target mixed noise dataset is constructed according to the noise data obtained after processing. Specifically, the target devices need to simulate the device working state corresponding to a plurality of load conditions and fault states in the target environment, and the device noise generated by the target devices in the device working state is collected through a pre-set microphone array. It needs to be noted that the object of noise collection is the sound-emitting power transmission and transformation equipment such as the main transformer and high-voltage shunt reactor used in 110kV to 500kV substations, so the target devices also simulate the noise generated in this load condition and the corresponding fault state. It needs to be further noted that the target environment refers to an anechoic chamber environment, and the device refers to a combination system of a gun-shaped electret condenser microphone and a recorder, thereby ensuring the reliability of the recorded data, and the microphone is about 4m away from the device. By adjusting the load condition of the device and simulating different fault states, noise data of various sound pressure levels and frequency spectrum distributions can be generated; wherein each segment of the collected noise of a single device running is taken as a target output and the mixed noise of multiple devices running simultaneously is taken as an input to form a group of samples. The device noise includes pure noise generated by a single target device in a plurality of target devices and noise generated by no less than one target device.
[0061] Further, the collected device noise needs to be preprocessed. Specifically, noise data with a time length less than a pre-set time length threshold needs to be removed from the device noise to obtain removed noise data, and the removed noise data is subjected to Hilbert transform to obtain a complex time domain analytic signal. It needs to be noted that the device noise dataset in the substation scenario is constructed by collecting the background noise in the real substation scenario and fusing it with the multi-channel audio data of the electrical equipment to generate station mixed noise samples with a random signal-to-noise ratio of -5 to 5dB. Then, audio segments with a time length less than 3 seconds in the samples are discarded, and the multi-channel noisy mixed signal is subjected to Hilbert transform to obtain a complex time domain analytic signal containing real and imaginary parts. Then, the target mixed noise dataset is constructed according to the complex time domain analytic signal.
[0062] Step S12, mixing the original audio signals to obtain a mixed audio signal, decomposing the mixed audio signal by a preset sound separation model to obtain a decomposed audio signal, and constructing a signal distortion ratio and a scale-invariant signal distortion ratio by using the original audio signal, the decomposed audio signal and the mixed audio signal.
[0063] In the embodiment, to construct the signal distortion ratio and the scale-invariant signal distortion ratio, first, two clean original audio signals and are mixed to obtain a mixed audio signal , which is denoted by s in the following and , and the mixed audio signal is decomposed by a preset sound separation model to obtain a decomposed audio signal , , which is denoted by in the following , .
[0064] To construct the signal distortion ratio SDRi (Signal-to-Distortion Ratio Improvement), a difference between the original audio signal s and the mixed audio signal is determined to obtain a first difference , and a difference between the original audio signal s and the decomposed audio signal is determined to obtain a second difference , then a ratio between a square of a norm of the mixed audio signal and a square of a norm of the first difference is determined to obtain a first ratio , and a ratio between a square of a norm of the original audio signal and a square of a norm of the second difference is determined to obtain a second ratio , a ten-base logarithm operation of ten times of the first ratio is performed to obtain a first operation result, that is , and a ten-base logarithm operation of ten times of the second ratio is performed to obtain a second operation result, that is , and a difference between the first operation result and the second operation result is taken as the signal distortion ratio, that is .
[0065] To construct the scale-invariant signal distortion ratio SI-SDR (Scale-invariant Signal-to-Distortion Ratio), a difference between the mixed audio signal The inner product of the original audio signal s is used to obtain the first inner product. And determine the ratio between the first inner product and the square of the norm of the original audio signal, so as to use the obtained ratio as the first scalar. Then, the product of the first scalar and the original audio signal is used as the first processed audio signal, that is... And calculate the difference between the mixed audio signal and the first processed audio signal to obtain the first signal difference, that is... It is also necessary to calculate the square of the norm of the first processed audio signal. Square of the norm of the difference from the first signal The ratio between them, to obtain the third ratio. And determine the inner product of the decomposed audio signal and the original audio signal to obtain the second inner product. And determine the ratio between the second inner product and the square of the norm of the original audio signal, so that the obtained ratio can be used as the second scalar, that is... The product of the second scalar and the original audio signal is used as the second processed audio signal. The difference between the decomposed audio signal and the second processed audio signal is calculated to obtain the second signal difference, i.e. The fourth ratio is obtained by calculating the ratio between the square of the norm of the second processed audio signal and the square of the norm of the difference between the two signals. Then, perform a logarithmic operation with base 10 on the third ratio to obtain the third result, which is... Then, perform a logarithmic operation with base 10 on the fourth ratio to obtain the fourth result. Finally, the difference between the third and fourth operation results is used as the scale-invariant signal distortion ratio, that is... .
[0066] Step S13: Optimize the preset multi-channel speech enhancement model using the signal distortion ratio, the scale-invariant signal distortion ratio, and the target mixed noise dataset to obtain the target multi-channel speech enhancement model.
[0067] In this embodiment, a loss function for a preset multi-channel speech enhancement model needs to be constructed based on the signal-to-distortion ratio. Specifically, the negative of the signal-to-distortion ratio needs to be used as the loss function for the preset multi-channel speech enhancement model. Then, the preset multi-channel speech enhancement model is trained using a target mixed noise dataset to obtain the multi-channel speech enhancement model to be validated. Specifically, the target mixed noise dataset needs to be divided according to a preset ratio, for example, into a training set T1, a validation set T2, and a test set T3 in a 10:1:1 ratio. The preset multi-channel speech enhancement model is then trained using the training set T1, and the optimal hyperparameter configuration of the network is optimized using a random search method on the validation set T2. It should be further noted that the preset multi-channel speech enhancement model is a model based on a complex adaptive beamforming network and a complex fully convolutional network, that is, the preset multi-channel speech enhancement model is CNABTCN (Complex Neural Network Adaptive Beamforming and Temporal Convolutional Network). Wherein, as... Figure 2 As shown, the CNABTCN network consists of a Complex-valued Adaptive Beamforming Network (CNAB) and a Complex Fully Convolutional Network (CFCN). It should be noted that a training environment needs to be set up before model training. In this environment, the number of training epochs is set to 100, with an initial learning rate of 0.0001 (the learning rate is halved if the validation set accuracy does not improve for three consecutive epochs), and the Adam optimizer is used.
[0068] Further, the performance of the to-be-confirmed multi-channel speech enhancement model needs to be verified by a signal distortion ratio and a scale-invariant signal distortion ratio. If the model performance of the to-be-confirmed multi-channel speech enhancement model reaches a preset performance condition, the to-be-confirmed multi-channel speech enhancement model is taken as the target multi-channel speech enhancement model. If the model performance of the to-be-confirmed multi-channel speech enhancement model does not reach the preset performance condition, the step of training the preset multi-channel speech enhancement model by the target mixed noise data set is jumped to, to obtain the to-be-confirmed multi-channel speech enhancement model, until the model performance of the to-be-confirmed multi-channel speech enhancement model reaches the preset performance condition. The signal distortion ratio SDRi measures the improvement of signal quality by calculating the SDR difference between the separated audio and the pure device noise audio. The scale-invariant signal distortion ratio SI-SDRi evaluates the separation accuracy by calculating the direction proximity between the separated audio and the pure device noise audio. The performance verification needs to be verified by the test set T3, and the preset performance condition can be adjusted according to the requirement. In this way, the model performance can be guaranteed, and the accuracy of noise separation is improved.
[0069] In step S14, noise separation is performed on the to-be-processed noise data based on the target multi-channel speech enhancement model, to obtain a separated noise signal and a separated audio signal.
[0070] In this embodiment, the to-be-processed noise needs to be separated by the target multi-channel speech enhancement model trained. Specifically, the to-be-processed noise needs to be convolved by the complex adaptive beamforming network in the target multi-channel speech enhancement model, to obtain an enhanced noise signal. It needs to be noted that the complex adaptive beamforming network CNAB is mainly composed of a complex LSTM (Long Short-Term Memory) and a complex linear layer linear, which is used to estimate the beamforming filter coefficients, such as Figure 3 The rules of the complex LSTM are provided in the above formula, wherein the first layer of the CNAB is a complex LSTM, which is called a complex shared LSTM, which takes the complex time-domain waveform in the target mixed noise data set as input. The next layer has two separated complex LSTMs, which correspond to processing the features in the corresponding channels, which are called complex separated LSTMs. Finally, the beamforming filter is generated through a complex linear activation layer, and the enhanced speech features are estimated through a complex convolution operation. The specific operation mode of the complex LSTM is as follows:
[0071] ;
[0072] ;
[0073] ;
[0074] wherein, and is composed of two normal LSTM layers, which represent the real part and the imaginary part of the complex LSTM, respectively, and are the real part and the imaginary part of the feature map. is the complex sequence of the complex LSTM output feature map. Wherein, is the complex analytic signal after Hilbert transform, a represents that it is a complex analytic signal, and c is the channel index of the microphone array.
[0075] Further, the operation mode of the complex linear layer is similar to that of the complex LSTM layer, and the output of the complex linear layer is the complex beamforming filter coefficient estimated by the CNAB module . Finally, the output of the CNAB module is:
[0076] ;
[0077] wherein, represents a complex convolution operation, which sums the complex convolution structures corresponding to two channels to obtain the output of the complex adaptive beamforming network. It should be noted that is a single-channel complex signal.
[0078] Further, the time-domain complex ratio mask of the enhanced noise signal needs to be calculated by the complex fully convolutional network, and the noise to be processed is noise separated by the time-domain complex ratio mask to obtain the separated noise signal and the separated audio signal. Wherein, for the complex fully convolutional network CFCN part, it is composed of Encoder, local complex TCN module and Decoder. The Encoder is composed of one one-dimensional convolution layer, the output of the CNAB is taken as the input quantity of the Encoder, and the weight sharing is passed, as follows:
[0079] ;
[0080] ;
[0081] wherein, represents the feature mapping function of the Encoder layer, and represent the real part and the imaginary part features mapped by the Encoder layer, respectively.
[0082] Wherein, the local complex TCN network module is composed of layer normalization, complex one-dimensional convolution and local complex TCN network. The role of layer normalization is to speed up network training and prevent gradient explosion or gradient disappearance. The intermediate features output by the Encoder are normalized in the following way:
[0083] ;
[0084] wherein is the real part of the after processing, is the feature mapping function of the normalization layer, is the normalized feature. The complex one-dimensional convolution layer processes the input and as follows:
[0085] ;
[0086] From the above, it can be seen that the complex one-dimensional convolution is composed of two ordinary one-dimensional convolution layers, wherein corresponds to the feature mapping function of the real convolution layer, represents the feature mapping corresponding to the virtual convolution layer, is the final output of the complex one-dimensional convolution layer. A local complex TCN network is connected after the complex one-dimensional convolution layer, as shown in Figure 4 . In order to make full use of the time context window of the speech signal, the dilation factor of the 1-D Conv block in each layer is exponentially increasing. In the local complex TCN network, there are X 1-D Convs, and the dilation factors are 1, 2, 4, …, , repeated R times. As shown in Figure 4 , different colors are used to represent 1-D Convs with different dilation factors, wherein gray is a complex 1-D Conv block.
[0087] The local complex TCN network estimates the time-domain complex ratio mask CRM, given the clean speech time-domain analytic signal s and the time-domain noisy analytic signal x, and the CRM is defined as:
[0088] ;
[0089] wherein and are the real part and the imaginary part of the time-domain noisy analytic signal, and are the real part and the imaginary part of the analytic signal. Finally, the CRM can be input into the Decoder network, and the model estimates the analytic signal by minimizing the loss function to approximate the target speech analytic signal. Finally, the real part of the analytic signal output by the model is taken as the separated audio signal, and the imaginary part of the analytic signal is taken as the separated noise signal.
[0090] In the embodiment, device noises generated when several target devices in a target environment work can be collected, and the device noises are preprocessed to construct a target mixed noise dataset based on obtained preprocessed noise data; the target mixed noise dataset is a dataset simulating substation noise; an original audio signal is mixed to obtain a mixed audio signal, and the mixed audio signal is decomposed by a preset sound separation model to obtain a decomposed audio signal, and then a signal distortion ratio and a scale-invariant signal distortion ratio are constructed by the original audio signal, the decomposed audio signal, and the mixed audio signal; a preset multi-channel speech enhancement model is optimized by the signal distortion ratio, the scale-invariant signal distortion ratio, and the target mixed noise dataset to obtain a target multi-channel speech enhancement model; and noise separation is performed on to-be-processed noise data based on the target multi-channel speech enhancement model to obtain a separated noise signal and a separated audio signal. As can be seen, the method of the application can collect device noises generated when target devices in a target environment work, and then preprocess the device noises to construct a target mixed noise dataset based on obtained preprocessed noise data, and then mix an original audio signal, decompose the obtained mixed audio signal to obtain a decomposed audio signal, and then construct a signal distortion ratio and a scale-invariant signal distortion ratio by the original audio signal, the decomposed audio signal, and the mixed audio signal. The preset multi-channel speech enhancement model is optimized according to the distortion ratio and the target mixed noise dataset, and noise separation is performed on the target mixed noise dataset based on the obtained target multi-channel speech enhancement model to obtain a separated noise signal and a separated audio signal. In this way, the noise signal can be processed by the improved multi-channel speech enhancement model, and the effect of noise separation is improved.
[0091] Referring to Figure 5 The embodiment of the application discloses a substation noise intelligent separation device, which comprises:
[0092] A dataset construction module 11 is configured to collect device noises generated when several target devices in a target environment work, and preprocess the device noises to construct a target mixed noise dataset based on obtained preprocessed noise data; the target mixed noise dataset is a dataset simulating substation noise;
[0093] A distortion parameter calculation module 12 is configured to mix an original audio signal to obtain a mixed audio signal, decompose the mixed audio signal by a preset sound separation model to obtain a decomposed audio signal, and then construct a signal distortion ratio and a scale-invariant signal distortion ratio by the original audio signal, the decomposed audio signal, and the mixed audio signal;
[0094] The model optimization module 13 is configured to optimize a preset multi-channel speech enhancement model based on the signal distortion ratio, the scale-invariant signal distortion ratio, and the target mixed noise data set, to obtain a target multi-channel speech enhancement model.
[0095] The noise separation module 14 is configured to perform noise separation on to-be-processed noise data based on the target multi-channel speech enhancement model, to obtain a separated noise signal and a separated audio signal.
[0096] In this embodiment, device noise generated by a plurality of target devices in a target environment can be collected, and the device noise can be preprocessed to construct a target mixed noise data set based on the obtained preprocessed noise data; the target mixed noise data set is a data set simulating substation noise; an original audio signal can be mixed to obtain a mixed audio signal, and the mixed audio signal can be decomposed by a preset sound separation model to obtain a decomposed audio signal, and then the signal distortion ratio and the scale-invariant signal distortion ratio can be constructed based on the original audio signal, the decomposed audio signal, and the mixed audio signal; a preset multi-channel speech enhancement model can be optimized based on the signal distortion ratio, the scale-invariant signal distortion ratio, and the target mixed noise data set, to obtain a target multi-channel speech enhancement model; and to-be-processed noise data can be separated based on the target multi-channel speech enhancement model, to obtain a separated noise signal and a separated audio signal. As can be seen, the method of the present application can collect device noise generated by target devices in a target environment, and then preprocess the device noise to construct a target mixed noise data set based on the obtained preprocessed noise data. Then, the original audio signal needs to be mixed, and the obtained mixed audio signal needs to be decomposed to obtain a decomposed audio signal. Then, the signal distortion ratio and the scale-invariant signal distortion ratio are constructed based on the original audio signal, the decomposed audio signal, and the mixed audio signal. And the preset multi-channel speech enhancement model needs to be optimized based on the above distortion ratio and the target mixed noise data set, to separate the target mixed noise data set based on the obtained target multi-channel speech enhancement model, to obtain a separated noise signal and a separated audio signal. In this way, the noise signal can be processed by the improved multi-channel speech enhancement model, and the effect of noise separation is improved.
[0097] In some embodiments, the data set construction module 11 can specifically include:
[0098] The noise collection unit is configured to simulate a plurality of load conditions and corresponding device working states of a plurality of target devices in a target environment, and collect device noise generated by the target devices in the device working states through a preset microphone array; the device noise includes noise generated by a single target device and noise generated by no less than one target device;
[0099] The data elimination unit is configured to eliminate noise data with a time length less than a preset time length threshold in the device noise to obtain post-elimination noise data, and perform Hilbert transform on the post-elimination noise data and a preset background noise signal to obtain a complex time domain analytic signal; the preset background noise signal is a pre-collected substation background noise;
[0100] The data set construction unit is configured to construct a target mixed noise data set based on the complex time domain analytic signal.
[0101] In some embodiments, the distortion parameter calculation module 12 can specifically include:
[0102] The first difference calculation unit is configured to determine a difference between the original audio signal and the mixed audio signal to obtain a first difference, and determine a difference between the original audio signal and the post-decomposition audio signal to obtain a second difference;
[0103] The first ratio calculation unit is configured to determine a ratio between a square of a norm of the mixed audio signal and a square of a norm of the first difference to obtain a first ratio, and determine a ratio between a square of a norm of the original audio signal and a square of a norm of the second difference to obtain a second ratio;
[0104] The first logarithm operation unit is configured to perform ten times base-ten logarithm operation on the first ratio to obtain a first operation result, and perform ten times base-ten logarithm operation on the second ratio to obtain a second operation result;
[0105] The signal distortion ratio determination unit is configured to take a difference between the first operation result and the second operation result as a signal distortion ratio.
[0106] In some embodiments, the distortion parameter calculation module 12 can specifically include:
[0107] The first scalar determination unit is configured to determine an inner product of the mixed audio signal and the original audio signal to obtain a first inner product, and determine a ratio between the first inner product and a square of a norm of the original audio signal to obtain a first scalar;
[0108] a second difference calculation unit configured to calculate a difference between the mixed audio signal and a first processed audio signal which is a product of the first scalar and the original audio signal, to obtain a first signal difference;
[0109] a second ratio calculation unit configured to calculate a ratio between a square of a norm of the first processed audio signal and a square of a norm of the first signal difference, to obtain a third ratio;
[0110] a second scalar determination unit configured to determine an inner product of the decomposed audio signal and the original audio signal to obtain a second inner product, and determine a ratio between the second inner product and a square of a norm of the original audio signal, to obtain a second scalar;
[0111] a third difference calculation unit configured to calculate a difference between the second scalar and a second processed audio signal which is a product of the second scalar and the original audio signal, to obtain a second signal difference;
[0112] a third ratio calculation unit configured to calculate a ratio between a square of a norm of the second processed audio signal and a square of a norm of the second signal difference, to obtain a fourth ratio;
[0113] a second logarithm operation unit configured to perform a ten-base logarithm operation on the third ratio by ten to obtain a third operation result, and perform a ten-base logarithm operation on the fourth ratio by ten to obtain a fourth operation result;
[0114] a scale-invariant signal-to-distortion ratio determination unit configured to determine a scale-invariant signal-to-distortion ratio based on a difference between the third operation result and the fourth operation result.
[0115] In some embodiments, the model optimization module 13 can specifically include:
[0116] a loss function construction unit configured to construct a loss function of the preset multi-channel speech enhancement model based on the signal-to-distortion ratio;
[0117] a model training unit configured to train a preset multi-channel speech enhancement model based on the target mixed noise data set, to obtain a to-be-confirmed multi-channel speech enhancement model;
[0118] a performance verification unit configured to perform performance verification on the to-be-confirmed multi-channel speech enhancement model based on the signal-to-distortion ratio and the scale-invariant signal-to-distortion ratio;
[0119] The model confirmation unit is used to select the multi-channel speech enhancement model to be confirmed as the target multi-channel speech enhancement model if the model performance of the multi-channel speech enhancement model to be confirmed meets the preset performance conditions.
[0120] The step jump unit is used to jump to the step of training the preset multi-channel speech enhancement model with the target mixed noise dataset to obtain the multi-channel speech enhancement model to be confirmed if the model performance of the multi-channel speech enhancement model to be confirmed does not meet the preset performance conditions, until the model performance of the multi-channel speech enhancement model to be confirmed reaches the preset performance conditions.
[0121] In some embodiments, the preset multi-channel speech enhancement model is a model based on a complex adaptive beamforming network and a complex fully convolutional network.
[0122] In some embodiments, the noise separation module 14 may specifically include:
[0123] The signal enhancement unit is used to perform a convolution operation on the noise to be processed through the complex adaptive beamforming network in the target multi-channel speech enhancement model to obtain the enhanced noise signal.
[0124] The noise separation unit is used to calculate the temporal complex ratio mask of the enhanced noise signal through the complex fully convolutional network, and to separate the noise to be processed through the temporal complex ratio mask to obtain the separated noise signal and the separated audio signal.
[0125] Furthermore, embodiments of this application also disclose an electronic device, Figure 6 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.
[0126] Figure 6 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the substation noise intelligent separation method disclosed in any of the foregoing embodiments. Alternatively, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0127] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.
[0128] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0129] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the substation noise intelligent separation method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs capable of performing other specific tasks.
[0130] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned intelligent substation noise separation method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.
[0131] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.
[0132] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0133] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0134] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0135] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for intelligent separation of noise in substations, characterized in that, include: The noise generated by several target devices in the target environment during operation is collected, and the device noise is preprocessed to construct a target mixed noise dataset based on the obtained preprocessed noise data; the target mixed noise dataset is a dataset simulating substation noise. The original audio signal is mixed to obtain a mixed audio signal, and the mixed audio signal is decomposed by a preset sound separation model to obtain a decomposed audio signal. Then, the signal distortion ratio and the scale-invariant signal distortion ratio are constructed by the original audio signal, the decomposed audio signal and the mixed audio signal. The preset multi-channel speech enhancement model is optimized using the signal distortion ratio, the scale-invariant signal distortion ratio, and the target mixed noise dataset to obtain the target multi-channel speech enhancement model. Based on the target multi-channel speech enhancement model, noise separation is performed on the noise data to be processed to obtain the separated noise signal and the separated audio signal.
2. The intelligent noise separation method for substations according to claim 1, characterized in that, The method involves collecting device noise generated by several target devices operating in the target environment, preprocessing the device noise, and constructing a target mixed noise dataset based on the preprocessed noise data, including: In a target environment, several load conditions and fault states are simulated using several target devices, and the device noise generated by the target devices under these operating states is collected using a preset microphone array; the device noise includes the noise generated by a single target device and the noise generated by at least one target device. Noise data with a time length less than a preset time length threshold is removed from the equipment noise to obtain noise data after removal. Hilbert transform is then performed on the noise data after removal and a preset background noise signal to obtain a complex time-domain analytic signal. The preset background noise signal is the pre-collected substation background noise. A target mixed noise dataset is constructed based on the complex time-domain analytic signal.
3. The intelligent noise separation method for substations according to claim 1, characterized in that, Constructing the signal distortion ratio using the original audio signal, the decomposed audio signal, and the mixed audio signal includes: The difference between the original audio signal and the mixed audio signal is determined to obtain a first difference, and the difference between the original audio signal and the decomposed audio signal is determined to obtain a second difference; Determine the ratio between the square of the norm of the mixed audio signal and the square of the norm of the first difference to obtain a first ratio, and determine the ratio between the square of the norm of the original audio signal and the square of the norm of the second difference to obtain a second ratio. Perform a logarithmic operation with the base 10 on the first ratio to obtain a first result, and perform a logarithmic operation with the base 10 on the second ratio to obtain a second result. The difference between the first calculation result and the second calculation result is used as the signal distortion ratio.
4. The intelligent noise separation method for substations according to claim 1, characterized in that, Constructing a scale-invariant signal distortion ratio using the original audio signal, the decomposed audio signal, and the mixed audio signal includes: Determine the inner product of the mixed audio signal and the original audio signal to obtain a first inner product, and determine the ratio between the first inner product and the square of the norm of the original audio signal, so as to use the obtained ratio as a first scalar; The product of the first scalar and the original audio signal is used as the first processed audio signal, and the difference between the mixed audio signal and the first processed audio signal is calculated to obtain the first signal difference. Calculate the ratio between the square of the norm of the first processed audio signal and the square of the norm of the first signal difference to obtain a third ratio; The inner product of the decomposed audio signal and the original audio signal is determined to obtain a second inner product, and the ratio between the second inner product and the square of the norm of the original audio signal is determined to be used as a second scalar. The product of the second scalar and the original audio signal is used as the second processed audio signal, and the difference between the decomposed audio signal and the second processed audio signal is calculated to obtain the second signal difference. Calculate the ratio between the square of the norm of the second processed audio signal and the square of the norm of the second signal difference to obtain the fourth ratio; Perform a logarithmic operation with base 10 on the third ratio to obtain the third result, and perform a logarithmic operation with base 10 on the fourth ratio to obtain the fourth result. The difference between the third calculation result and the fourth calculation result is used as the scale-invariant signal distortion ratio.
5. The intelligent noise separation method for substations according to claim 1, characterized in that, The optimization of the preset multi-channel speech enhancement model using the signal distortion ratio, the scale-invariant signal distortion ratio, and the target mixed noise dataset to obtain the target multi-channel speech enhancement model includes: The loss function of the preset multi-channel speech enhancement model is constructed based on the signal distortion ratio; The preset multi-channel speech enhancement model is trained using the target mixed noise dataset to obtain the multi-channel speech enhancement model to be confirmed. The performance of the multi-channel speech enhancement model to be confirmed is verified based on the signal distortion ratio and the scale-invariant signal distortion ratio. If the performance of the multi-channel speech enhancement model to be confirmed reaches the preset performance condition, then the multi-channel speech enhancement model to be confirmed will be used as the target multi-channel speech enhancement model. If the performance of the multi-channel speech enhancement model to be confirmed does not meet the preset performance conditions, the process jumps to the step of training the preset multi-channel speech enhancement model using the target mixed noise dataset to obtain the multi-channel speech enhancement model to be confirmed, until the performance of the multi-channel speech enhancement model to be confirmed reaches the preset performance conditions.
6. The intelligent noise separation method for substations according to any one of claims 1 to 5, characterized in that, The preset multi-channel speech enhancement model is a model based on a complex adaptive beamforming network and a complex fully convolutional network.
7. The intelligent noise separation method for substations according to claim 6, characterized in that, The noise separation based on the target multi-channel speech enhancement model to obtain the separated noise signal and the separated audio signal includes: The complex adaptive beamforming network in the target multi-channel speech enhancement model performs a convolution operation on the noise to be processed to obtain the enhanced noise signal. The enhanced noise signal is calculated using the complex full convolutional network, and the noise to be processed is separated using the complex ratio mask to obtain the separated noise signal and the separated audio signal.
8. A substation noise intelligent separation device, characterized in that, include: The dataset construction module is used to collect equipment noise generated by several target devices in the target environment during operation, and to preprocess the equipment noise to construct a target mixed noise dataset based on the obtained preprocessed noise data; the target mixed noise dataset is a dataset simulating substation noise. The distortion parameter calculation module is used to mix the original audio signal to obtain a mixed audio signal, and decompose the mixed audio signal through a preset sound separation model to obtain a decomposed audio signal. Then, the signal distortion ratio and the scale-invariant signal distortion ratio are constructed by the original audio signal, the decomposed audio signal and the mixed audio signal. The model optimization module is used to optimize the preset multi-channel speech enhancement model using the signal distortion ratio, the scale-invariant signal distortion ratio, and the target mixed noise dataset to obtain the target multi-channel speech enhancement model. The noise separation module is used to separate noise from the noise data to be processed based on the target multi-channel speech enhancement model, so as to obtain the separated noise signal and the separated audio signal.
9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the substation noise intelligent separation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the intelligent substation noise separation method as described in any one of claims 1 to 7.