Method and device for identifying opening and closing sounds of earthing switches
By combining spectrum analysis and deep learning technology, identifying the sound characteristics of the grounding knife switch, the problem of low sound recognition accuracy in the prior art is solved, and the accuracy and safety of robot operation are achieved.
Patent Information
- Application Number
- CN202111513750.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-12-13
AI Technical Summary
In the prior art, when using robots to perform grounding knife switch opening and closing operations, the sound recognition accuracy is not high and is easily disturbed by noise, resulting in operation failure and safety hazards.
The methods of real-time data acquisition, data analysis and conversion, feature extraction, identification parameter acquisition and convolutional neural network recognition are adopted. Through a combination of spectrum analysis and deep learning, the sound characteristics of the grounding knife switch are identified, other noises are eliminated, and the recognition accuracy is improved.
It realizes rapid and accurate identification of the sound of the grounding knife switch opening and closing, ensuring the success rate of robot operation and reducing safety risks.
Smart Images

Figure CN114283792B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sound recognition, and particularly relates to a method and device for recognizing the closing and opening sounds of an earthing switch. Background Art
[0002] When performing maintenance on power equipment, it is necessary to perform closing and opening operations on the earthing switch. The sound of the switch is a key factor in judging whether the closing and opening operations of the switch are successful.
[0003] Currently, when using a robot to operate the earthing switch in the switchgear room, the sound decibel level generated during the closing and opening of the switch is often used to judge whether the operation is successful. However, there are many other noises during the operation, such as the sound of people talking, the sound of car horns, abnormal noises of motors, and the sound of the robot contacting the cabinet. Relying solely on the sound decibel will result in many misidentifications, leading to operation failures, affecting the success rate of robot operations, and posing a great potential safety hazard.
[0004] In summary, the problem of the existing technology is that the recognition accuracy of the closing and opening sounds of the earthing switch is not high. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for recognizing the closing and opening sounds of an earthing switch with high accuracy.
[0006] Another purpose of the present invention is to provide a device for recognizing the closing and opening sounds of an earthing switch with high accuracy.
[0007] The technical solution for achieving the purpose of the present invention is as follows:
[0008] A method for recognizing the closing and opening sounds of an earthing switch includes the following steps:
[0009] (10) Real-time data acquisition: Obtain dual-channel real-time sound data through a sound card module;
[0010] (20) Data parsing and conversion: Decode the mono sound data extracted from the dual-channel real-time sound data to obtain audio data;
[0011] (30) Feature extraction: Perform preprocessing and fast Fourier transform on the audio data, calculate the current voice data energy, obtain the energy change situation at different frequencies and time domains, and perform high-frequency and medium-frequency filtering based on this energy to extract the audio characteristics of the signal;
[0012] (40) Recognition parameter acquisition: Filter other sounds based on the audio characteristics of the signal to obtain the recognition parameters of the earthing switch sound characteristics;
[0013] (50) Sound recognition: Based on a convolutional neural network, the on-site sound data is recognized by means of a sliding window to obtain the recognition result of the grounding switch sound.
[0014] (60) Output of recognition result: Combine the characteristic recognition parameters and the recognition result of the grounding switch sound obtained based on the convolutional neural network.
[0015] The technical solution for achieving another object of the present invention is as follows:
[0016] A grounding switch opening and closing sound recognition device includes:
[0017] A real-time data acquisition module for acquiring dual-channel real-time sound data through a sound card module;
[0018] A data parsing and conversion module for decoding the mono sound data extracted from the dual-channel real-time sound data to obtain audio data;
[0019] A feature extraction module for preprocessing the audio data and performing a fast Fourier transform, calculating the current speech data energy, obtaining the energy change situation under different frequencies and time domains, and performing high-frequency and medium-frequency filtering based on this energy to extract the audio characteristics of this signal;
[0020] A recognition parameter acquisition module for filtering other sounds based on the audio characteristics of this signal to obtain the grounding switch sound characteristic recognition parameters;
[0021] A sound recognition module for recognizing the on-site sound data by means of a sliding window based on a convolutional neural network to obtain the recognition result of the grounding switch sound.
[0022] A recognition result output module for combining the characteristic recognition parameters and the recognition result of the grounding switch sound obtained based on the convolutional neural network. Compared with the prior art, the present invention has the following remarkable advantages:
[0023] High accuracy: The method of the present invention comprehensively uses the spectrum analysis method and the deep learning method to recognize sounds, achieving the purpose of quickly and accurately recognizing the opening and closing sounds of the grounding switch and excluding other noises, ensuring the successful operation of the robot grounding knife.
[0024] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Description of the Drawings
[0025] Figure 1 is the main flowchart of the grounding switch opening and closing sound recognition method of the present invention.
[0026] Figure 2 is Figure 1 the flowchart of the feature extraction step in
[0027] Figure 3 It is a flowchart of the frame division process.
[0028] Figure 4 It is a flowchart of the windowing implementation process.
[0029] Figure 5 It is Figure 1 a flowchart of the step for obtaining recognition parameters in
[0030] Figure 6 It is a convolutional neural model diagram.
[0031] Figure 7 It is an example diagram of a residual network including Bottleneck.
[0032] Figure 8 It is a schematic diagram of sound recognition using a sliding window. Specific implementation manner
[0033] As Figure 1 shown, the method for recognizing the opening and closing sounds of the grounding switch of the present invention includes the following steps:
[0034] (10) Real-time data acquisition: Obtain dual-channel real-time sound data through a sound card module;
[0035] The real-time sound data can be obtained through an internal or external sound card module, and after A / D analog-to-digital conversion, it is encoded after 16-bit, dual-channel, and 48K sampling;
[0036] (20) Data parsing and conversion: Decode the mono sound data extracted from the dual-channel real-time sound data to obtain audio data;
[0037] The process of data parsing and conversion is specifically as follows: First, the obtained real-time data is dual-channel data, and the data processing process only needs to analyze the data of a single channel. Therefore, it is necessary to extract mono data from the dual-channel data, and then decode the mono data to obtain the audio data to be processed. The process of data decoding is to restore the data to the sampled digital signal according to the number of sampling bits and the big-end or little-end mode of the data;
[0038] (30) Feature extraction: Perform preprocessing and fast Fourier transform on the audio data, calculate the energy of the current voice data, obtain the energy change situation under different frequencies and time domains, and perform high-frequency and medium-frequency filtering based on this energy to extract the audio characteristics of the signal;
[0039] Feature extraction is performed on the signal. The main process is to first preprocess the data, which includes framing and windowing the data to make the original non-stationary signal stationary. Finally, a fast Fourier transform is performed on the preprocessed signal, and then the energy of the current speech data is calculated to obtain the energy change in different frequencies and time domains. Based on this energy, high-frequency and intermediate-frequency filtering are performed to extract the audio characteristics of the signal.
[0040] Figure 2 It is a flowchart of the feature extraction step. The main features include the different energy performance characteristics of the signal in the time domain and frequency domain. The present invention realizes the recognition and feature extraction of the signal based on the distribution of the energy strength of the signal in different frequency bands.
[0041] Such as Figure 2 shown, the (30) feature extraction step includes:
[0042] (31) Data preprocessing part: Frame and window the decoded audio data.
[0043] The processing process of the signal requires the data to be a stationary signal and meet certain periodicity requirements. It is necessary to preprocess the signal. The data preprocessing part is to frame and window the data after decoding.
[0044] Figure 3 It is the flowchart of framing. The process of framing is to divide the original data into a certain number of signals according to the frame length and frame shift. First, set the frame length and frame shift lengths required for framing, then obtain the number of frames that the entire signal needs to be separated according to the set frame length and frame shift, and then obtain the data contained in each frame according to the frame shift. Finally, complete the framing of the original signal to obtain a certain number of framed data.
[0045] At the same time, in order to ensure the data quality and reduce the frequency domain leakage during data analysis, it is necessary to window the data. The process of windowing is to reduce the frequency domain leakage process and superimpose the data with the window function.
[0046] Figure 4 It is the windowing process of the data: First, it is necessary to select the window function, then set different window lengths according to the selected window function, and finally superimpose the set window function with each frame of data to reduce the data leakage. The present invention selects the Hanning window, which has strong main lobe and side lobe capabilities and better realizes the integrity of the data.
[0047] (32) Fourier transform: Transform the characteristics of the signal in the time domain to the frequency domain, change the characteristics of the signal in the time domain, and obtain the data of each frame of the signal after Fourier transform.
[0048] (33) Audio energy acquisition: Square the amplitude of the transformed data for each frame to obtain the energy value, with each energy corresponding to each frequency;
[0049] (34) Filtering algorithm identification: Filter the energy at different frequencies to obtain the energy distribution at different frequencies and acquire the characteristics of the sound signal.
[0050] (40) Identification parameter acquisition: Based on the audio characteristics of the signal, filter other sounds to obtain the identification parameters for the sound characteristics of the earthing switch.
[0051] As Figure 5 shown, the steps of the (40) identification parameter acquisition include:
[0052] (41) Data quantization: Square the data after Fourier transform, calculate the logarithm of the squared data, and use the result as the energy of the signal.
[0053] Squaring means squaring the Fourier transform of the original data signal:
[0054] Assume the energy of an energy signal s(t) is E, then its energy is: If the Fourier transform of this signal, i.e., the spectral density, is S(f), then
[0055] (42) Energy distribution acquisition: Based on the window length, the overall data size, and the sampling frequency, obtain the frequency corresponding to each energy to get the energy distribution of the signal.
[0056] The energy distribution is the frequency-domain information corresponding to each energy after Fourier transform.
[0057] (43) Energy characteristic acquisition: Based on the energy distribution, obtain the energy characteristics in different frequency bands.
[0058] This energy is based on the energy in each frequency band. Through the distribution of this energy at each frequency, the main characteristic manifestations of the earthing switch sound in the frequency domain can be obtained.
[0059] (44) High-frequency band-pass filtering: Set the high-frequency frequency range of the band-pass filter, then set the threshold of the energy within this frequency band, and then obtain the energy characteristics within this frequency band according to the threshold.
[0060] (45) Medium-frequency band-pass filtering: Set the medium-frequency frequency range of the band-pass filter, then set the threshold of the energy within this frequency band, and then obtain the energy characteristics within this frequency band according to the threshold.
[0061] (46) Parameter characteristic preservation: Preserve the energy characteristics of the grounding knife sound in the medium and high frequency bands, and use this energy characteristic as the identification parameter for the grounding knife sound characteristics.
[0062] (50) Sound recognition: Based on a convolutional neural network, combined with the grounding knife sound characteristic identification parameter, identify the on-site sound data in a sliding window manner to obtain the grounding knife sound recognition result.
[0063] The (50) sound recognition step includes:
[0064] (51) Sound sample collection: Collect grounding knife audio samples with noise interference in the on-site operation environment;
[0065] The training of the deep learning model depends on a large number of sample data. To improve the robustness of the model and avoid overfitting, a large number of grounding knife sounds with noise interference, as well as non-grounding knife sounds such as metal impact sounds and human voices, are collected in the on-site operation environment.
[0066] (52) Model training: Extract the logarithmic spectrogram from the grounding knife audio samples, randomly divide them into a training set and a test set, and use the training set to train the convolutional neural model;
[0067] All the collected audio samples are randomly divided into a training set and a test set (the training data uses not only the grounding knife opening and closing sound data, but also various similar sound data according to the previous description, so it is all audio samples), and are randomly divided into a training set and a test set, and the training set is used to train the model. The network model structure is as Figure 6 shown, where K is the convolutional kernel and S is the stride. A 1*1 convolutional network structure is used in the BottleNeck, and a non-linear excitation is added to the learning representation of the previous layer, which improves the network expression ability, reduces the amount of computation, and also reduces the number of parameters. As Figure 7 is the residual network containing BootleNeck. The first 1x1 convolution is used to reduce the dimension, and the second 1x1 convolution is used to increase the dimension.
[0068] y l =h(x l )+F(x l , w l )
[0069] x l+1 =f(y l )
[0070] where x l is the input of the l-th layer in the residual unit, w l is the weight of the residual unit, F(x l , w l) is a functional module of the residual unit. h(x l ) = W' l x, where W' l is a 1x1 convolution operation.
[0071] f is the activation function ReLu, which can effectively increase the non-linearity of the model. Its function is defined as:
[0072] f(x) = max(0, x)
[0073] The entire network structure receives sound data of [1, 24000] as the model input. After passing through the convolutional layer and the fully connected layer, the training result is finally obtained through the Softmax loss function.
[0074] (53) Sliding window recognition: Use the trained mature convolutional neural model to recognize the on-site sound data in a sliding window manner to obtain the recognition result of the grounding knife switch sound.
[0075] (60) Recognition result output: Combine the recognition result of the grounding knife switch sound obtained through the characteristic recognition parameters and the convolutional neural network.
[0076] The entire recognition method actually uses two types:
[0077] One is the recognition method based on the sound spectrum characteristics and energy, and the other is the method of directly training a model on the sound data through deep learning and then performing recognition. The recognition results of the two methods are unified as the result reference to obtain the final recognition result.
[0078] Since the recognition method based on sound feature analysis has a low missed detection rate and a high false detection rate, the recognition result based on sound characteristic analysis is used as the "coarse" recognition result, and the recognition result based on the convolutional neural network is used as the "fine" recognition result. When method 1 recognizes as "yes", method 2 is used for a second judgment, and finally the recognition result of the grounding knife switch opening and closing sound is output.
[0079] The present invention uses a sliding window form for sound recognition to ensure the real-time performance of the grounding knife sound recognition. The industrial control computer sound card acquires environmental sounds at a sampling frequency of 48000HZ. The deep learning model recognition reads the sound card data at a frequency of every 10ms and 480 bytes of data, and performs recognition in the way that the data enters the network every 500ms. Initially, since the sound card does not acquire enough data for 500ms, the present invention fills in the 10ms data that first enters the network, so that real-time recognition can be performed without waiting for the data length to reach 500ms. The subsequent 10ms data is filled in with 480ms of data in the same way, and so on. Such as Figure 8 。
[0080] By adopting the method of building a deep learning network for the recognition of the sound of earthing switches, the recognition effect with a maximum delay of 10 ms can be achieved. At the same time, combined with the recognition method based on the sound spectrum energy characteristics, the recognition effect with an accuracy rate of 98.6% and a recall rate of 99.23% is realized. This greatly guarantees the safety when using a robot to operate the cabinet.
[0081] The technical solution adopted by the present invention is a method for recognizing the sound of earthing switches based on deep learning, which is specifically implemented according to the following steps:
[0082] Step 1: Collect sound samples. The training of the deep learning model depends on a large number of sample data. In order to improve the robustness of the model and avoid overfitting, a large number of earthing switch sounds with noise interference, as well as non-earthing switch sounds such as metal impact sounds and human voices, are collected in the on-site operation environment. The logarithmic spectrogram will be extracted from the collected audio samples and randomly divided into a training set and a test set, and the model will be trained with the training set. The network model structure is as Figure 6 shown, where K is the convolution kernel and S is the stride. A 1*1 convolution network structure is used in the BottleNeck, adding a non-linear excitation to the learning representation of the previous layer, improving the network expression ability, reducing the amount of computation and also reducing the number of parameters. As Figure 7 is a residual network containing BootleNeck. The first 1x1 convolution is used to reduce the dimension, and the second 1x1 convolution is used to increase the dimension.
[0083] y l = h(x l ) + F(x l , w l )
[0084] X l+1 = f(y l )
[0085] where x l is the input of the l-th layer in the residual unit, w l is the weight of the residual unit, and F(x l , w l ) is the function module of the residual unit. h(x l ) = W' l x, where W' l is a 1x1 convolution operation.
[0086] f is the activation function ReLu, which can effectively increase the non-linearity of the model, and its function is defined as:
[0087] f(x) = max(0, x)
[0088] The entire network structure receives voice data in the range of [1, 24000] as the model input. After passing through the convolutional layer and the fully connected layer, the training result is finally obtained through the Softmax loss function.
[0089] Step 2: Use the trained model for earthing switch voice recognition. The present invention uses a sliding window form for voice recognition to ensure the real-time performance of earthing switch voice recognition. The industrial computer sound card acquires environmental sounds at a sampling frequency of 48000HZ. The deep learning model reads the sound card data at a frequency of 480 bytes of data every 10ms, and performs recognition in the way that data enters the network every 500ms. Initially, since the sound card does not acquire enough data for 500ms, the present invention completes the 10ms data that first enters the network, so that real-time recognition can be performed without waiting for the data length to reach 500ms. For the subsequent 10ms data, 480ms of data is completed in the same way, and so on. As Figure 8 。
[0090] By adopting the method of building a deep learning network for earthing switch voice recognition, the recognition effect with a maximum delay of 10ms can be achieved. At the same time, combined with the recognition method based on the voice spectrum energy feature, the recognition effect with an accuracy rate of 98.6% and a recall rate of 99.23% is achieved. It greatly guarantees the safety when using a robot to operate the cabinet.
[0091] The earthing switch opening and closing voice recognition device of the present invention includes:
[0092] A real-time data acquisition module, used to acquire dual-channel real-time voice data through the sound card module;
[0093] A data parsing and conversion module, used to decode the mono voice data extracted from the dual-channel real-time voice data to obtain audio data;
[0094] A feature extraction module, used to preprocess and perform a fast Fourier transform on the audio data, calculate the current voice data energy, obtain the energy change situation under different frequencies and time domains, and perform high-frequency and medium-frequency filtering based on this energy to extract the audio characteristics of this signal;
[0095] The feature extraction module includes:
[0096] A data preprocessing unit, used to perform frame division and windowing processing on the decoded audio data;
[0097] A Fourier transform unit, used to transform the characteristics of the signal in the time domain to the frequency domain, change the characteristics of the signal in the time domain, and obtain the data of each frame of the signal after Fourier transform;
[0098] An audio energy acquisition unit, which is used to calculate the square of the amplitude of each frame of the transformed data to obtain an energy value, and each energy corresponds to each frequency;
[0099] A filtering algorithm recognition unit, which is used to filter the energy at different frequencies to obtain the distribution of the energy at different frequencies and obtain the characteristics of the sound signal;
[0100] An identification parameter acquisition module, which is used to filter other sounds based on the audio characteristics of the signal to obtain the identification parameters of the grounding knife sound characteristics;
[0101] The identification parameter acquisition module includes:
[0102] A data quantization unit, which is used to square the data after Fourier transform, calculate the logarithm of the squared data, and use the result as the energy of the signal;
[0103] An energy distribution acquisition unit, which is used to obtain the frequency corresponding to each energy according to the window length, the entire data size, and the sampling frequency, and obtain the energy distribution of the signal;
[0104] An energy characteristic acquisition unit, which is used to obtain the energy characteristics in different frequency bands according to the energy distribution;
[0105] A high-frequency band-pass filtering unit, which is used to set the high-frequency frequency range of the band-pass filtering, then set the threshold of the energy in this frequency band, and then obtain the energy characteristics in this frequency band according to the threshold;
[0106] An intermediate-frequency band-pass filtering unit, which is used to set the intermediate-frequency frequency range of the band-pass filtering, then set the threshold of the energy in this frequency band, and then obtain the energy characteristics in this frequency band according to the threshold;
[0107] A parameter characteristic storage unit, which is used to store the energy characteristics of the grounding knife sound in the middle and high frequency bands, and use this energy characteristic as the identification parameter of the grounding knife sound characteristic;
[0108] A sound recognition module, which is used to recognize the on-site sound data in a sliding window manner based on a convolutional neural network to obtain the grounding knife sound recognition result.
[0109] The sound recognition module includes:
[0110] A sound sample acquisition unit, which is used to acquire grounding knife audio samples with noise interference in the on-site operation environment;
[0111] A model training unit, which is used to extract logarithmic spectrograms from the grounding knife audio samples, randomly divide them into a training set and a test set, and use the training set to train the convolutional neural model;
[0112] A sliding window recognition unit is used to recognize on-site sound data in a sliding window manner using a trained mature convolutional neural model to obtain the recognition result of the grounding switch sound.
[0113] A recognition result output module is used to combine the recognition result of the grounding switch sound obtained through characteristic recognition parameters and based on the convolutional neural network.
Claims
1. A method for identifying the opening and closing sounds of an earthing switch, characterized in that, It includes the following steps: (10) Real-time data acquisition: Obtain dual-channel real-time sound data through the sound card module; (20) Data parsing and conversion: Decode the mono sound data extracted from the dual-channel real-time sound data to obtain audio data; (30) Feature extraction: Preprocess the audio data and perform a fast Fourier transform, calculate the current audio data energy, obtain the energy change situation at different frequencies and time domains, perform high-frequency and medium-frequency filtering based on the audio data energy, and extract the audio characteristics of the sound signal; (40) Identification parameter acquisition: Based on the audio characteristics of the sound signal, filter other sounds to obtain the identification parameters for the sound characteristics of the earthing switch; (50) Sound recognition: Based on a convolutional neural network, identify the on-site sound data in a sliding window manner to obtain the recognition result of the earthing switch sound; (60) Output of recognition result: Combine the identification result of the earthing switch sound obtained through the characteristic identification parameters and based on the convolutional neural network.
2. The method for identifying the opening and closing sounds of an earthing switch according to claim 1, characterized in that, The (30) feature extraction step includes: (31) Data preprocessing: Perform frame division and windowing on the decoded audio data; (32) Fourier transform: Transform the characteristics of the signal in the time domain to the frequency domain, change the characteristics of the signal in the time domain, and obtain the data after Fourier transform for each frame of the signal; (33) Audio energy acquisition: Square the amplitude of the data after transformation for each frame to obtain the energy value, and each energy corresponds to each frequency; (34) Filtering algorithm identification; Filter the audio data energy at different frequencies to obtain the energy distribution at different frequencies and obtain the characteristics of the sound signal.
3. The method for identifying the opening and closing sounds of an earthing switch according to claim 2, characterized in that, The (40) identification parameter acquisition step includes: (41) Data quantization: Square the data after Fourier transform, calculate the logarithm of the squared data, and use the result as the energy of the sound signal; (42) Energy distribution acquisition: According to the window length, the entire data size, and the sampling frequency, obtain the frequency corresponding to each energy to obtain the energy distribution of the signal; (43) Energy characteristic acquisition: According to the energy distribution, obtain the energy characteristics in different frequency bands; (44) High-frequency band-pass filtering: Set the high-frequency frequency range of the band-pass filtering, then set the threshold of the energy in this frequency band, and then obtain the energy characteristics in this frequency band according to the threshold; (45) Medium-frequency band-pass filtering: Set the medium-frequency frequency range of the band-pass filtering, then set the threshold of the energy in this frequency band, and then obtain the energy characteristics in this frequency band according to the threshold; (46) Parameter characteristic preservation: Preserve the energy characteristics of the earthing switch sound in the medium and high frequency bands, and use this energy characteristic as the identification parameter for the earthing switch sound characteristics.
4. The method for identifying the opening and closing sounds of an earthing switch according to claim 3, characterized in that, The (50) sound recognition step includes: (51) Sound sample collection: Collect earthing knife audio samples with noise interference in the on-site operation environment; (52) Model training: Extract the logarithmic spectrogram from the earthing knife audio samples, randomly divide them into a training set and a test set, and use the training set to train the convolutional neural model; (53) Sliding window recognition: Use the trained mature convolutional neural model to recognize the on-site sound data in a sliding window manner to obtain the recognition result of the earthing switch sound.
5. An apparatus for identifying the opening and closing sounds of an earthing switch, characterized in that, It includes: A real-time data acquisition module for acquiring dual-channel real-time sound data through a sound card module; A data parsing and conversion module for decoding the mono sound data extracted from the dual-channel real-time sound data to obtain audio data; A feature extraction module for preprocessing and performing a fast Fourier transform on the audio data, calculating the energy of the current audio data, obtaining the energy change in different frequencies and time domains, performing high-frequency and intermediate-frequency filtering based on the audio data energy, and extracting the audio characteristics of the sound signal; A recognition parameter acquisition module for filtering other sounds based on the audio characteristics of the sound signal to obtain the recognition parameters of the earthing switch sound characteristics; A sound recognition module for recognizing the on-site sound data in a sliding window manner based on a convolutional neural network to obtain the recognition result of the earthing switch sound; A recognition result output module for combining the recognition result of the earthing switch sound obtained through the characteristic recognition parameters and based on the convolutional neural network.
6. The apparatus for identifying the opening and closing sounds of an earthing switch according to claim 5, characterized in that, The feature extraction module includes: A data preprocessing unit for performing frame division and windowing on the decoded audio data; A Fourier transform unit for transforming the characteristics of the signal in the time domain to the frequency domain, changing the characteristics of the signal in the time domain, and obtaining the data of each frame of the signal after Fourier transform; An audio energy acquisition unit for squaring the amplitude of the data of each frame after transformation to obtain the energy value, with each energy corresponding to each frequency; A filtering algorithm recognition unit for filtering the audio data energy at different frequencies to obtain the energy distribution at different frequencies and obtain the characteristics of the sound signal.
7. The grounding switch opening and closing sound recognition device according to claim 6, characterized in that, The recognition parameter acquisition module includes: A data quantization unit for squaring the data after Fourier transform, calculating the logarithm of the squared data, and taking the result as the energy of the sound signal; An energy distribution acquisition unit for obtaining the frequency corresponding to each energy according to the window length, the entire data size, and the sampling frequency, and obtaining the energy distribution of the sound signal; An energy characteristic acquisition unit for obtaining the energy characteristics in different frequency bands according to the energy distribution; A high-frequency band-pass filtering unit for setting the high-frequency frequency range of the band-pass filter, then setting the threshold of the energy in this frequency band, and then obtaining the energy characteristics in this frequency band according to the threshold; An intermediate-frequency band-pass filtering unit for setting the intermediate-frequency frequency range of the band-pass filter, then setting the threshold of the energy in this frequency band, and then obtaining the energy characteristics in this frequency band according to the threshold; A parameter characteristic storage unit for storing the energy characteristics of the earthing switch sound in the middle and high frequency bands, and taking this energy characteristic as the recognition parameter of the earthing switch sound characteristics.
8. The grounding switch opening and closing sound recognition device according to claim 7, characterized in that, The sound recognition module includes: A sound sample acquisition unit for acquiring the earthing knife audio sample with noise interference in the on-site operation environment; A model training unit, configured to extract log spectrograms from the earthing switch audio samples, randomly divide them into a training set and a test set, and use the training set to train a convolutional neural model; A sliding window recognition unit, configured to use the trained mature convolutional neural model to recognize on-site sound data in a sliding window manner to obtain an earthing switch sound recognition result.
Citation Information
Patent Citations
Method for identifying sound scenes based on CNN (convolutional neural network) and random forest classification
CN108231067A
Specific group identification method, electronic device and computer readable storage medium
CN109119069A