A fully mechanized coal mining face centralized control instruction voice recognition control system
By integrating voice acquisition, processing, and control modules into the remote control system of the fully mechanized mining face, and using wavelet packet decomposition and neural network technology to denoise and enhance voice data, the problem of voice recognition in complex noise environments is solved, and efficient and accurate equipment control is achieved.
Patent Information
- Application Number
- CN202310180524.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-28
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2043-02-28
AI Technical Summary
The existing remote control system for fully mechanized mining faces suffers from long response times, poor timeliness, and is prone to misoperation. Furthermore, conventional voice recognition technology struggles to accurately recognize voice commands in complex and noisy environments, increasing safety hazards and coal loss.
The system employs a voice acquisition, processing, and control system, including modules for monitoring, acquisition, enhancement, recognition, display, and control. It uses techniques such as wavelet packet decomposition, adaptive thresholding, low-rank sparse matrix decomposition, and phase compensation to denoise and enhance the voice data. It then utilizes a multi-layer neural network model for recognition, generates text data, and sends control commands.
It improves the timeliness and accuracy of remote control of equipment in fully mechanized mining faces, reduces accidents caused by misoperation, and enhances the convenience and safety of equipment control.
Smart Images

Figure CN116312517B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of speech recognition processing technology, and more specifically, relates to a speech recognition and control system for centralized control commands of fully mechanized mining faces. Background Technology
[0002] Automation in fully mechanized mining faces typically employs a production model that prioritizes automatic control at the mining face, supplemented by remote intervention control from a monitoring center. This model frees workers to work at the monitoring center, where they control the equipment via a control panel. However, existing remote control methods, primarily relying on human intervention at the control panel, are prone to long response times, poor timeliness, and are susceptible to errors, safety hazards, and coal loss. Therefore, researching voice recognition-based remote control technology for fully mechanized mining faces is essential to improve the timeliness of remote control and prevent accidents caused by misoperation.
[0003] However, the geological environment of the fully mechanized mining face is complex and special, with a large number of working devices concentrated in the relatively closed environment of the fully mechanized mining face. The voice data in this environment is often mixed with a lot of background noise that is difficult to process, which makes it impossible for conventional speech recognition technology to accurately recognize the voice data of the fully mechanized mining face. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a voice recognition and control system for centralized control commands in fully mechanized mining faces that performs a series of voice noise reduction and voice enhancement processes on the collected voice data in the complex and special environment of fully mechanized mining faces, thereby improving the accuracy of voice recognition in specific environments.
[0005] The objective of this invention is achieved through the following technical solution: a voice recognition and control system for centralized control commands in a fully mechanized mining face, comprising a voice acquisition system, a voice processing system, and a voice control system; the voice acquisition system includes a monitoring module and an acquisition module, the voice processing system includes an enhancement module and a recognition module, and the voice control system includes a display module and a control module.
[0006] The monitoring module is used to maintain low-power operation when the system is in a sleep state, wait for a wake-up command, and wake up other modules of the system after receiving the wake-up command.
[0007] The acquisition module is used to collect voice control commands issued by staff through handheld terminal devices, central control centers, or ground dispatch centers after receiving a wake-up command, and to convert the collected voice control commands into audio data that can be recognized by computers through analog-to-digital conversion.
[0008] The enhancement module is used to enhance the acquired audio data to obtain its frequency domain characteristics. The specific method of enhancement processing is as follows:
[0009] Step 1: Perform wavelet packet decomposition on the audio data to obtain N layers 2. N Frequency domain sub-signals of different frequency bands;
[0010] Step 2: Calculate the energy entropy of each sub-signal using its wavelet packet coefficients;
[0011] Step 3: Obtain the high-frequency sub-signal using an adaptive thresholding method;
[0012] Step 4: Perform low-rank and sparse matrix decomposition on the high-frequency sub-signals to obtain speech data and background noise data;
[0013] Step 5: Optimize the phase spectrum using a phase compensation function, and recombine the amplitude spectrum of the noisy high-frequency sub-signal with the optimized phase spectrum to generate frequency domain features of the enhanced speech signal.
[0014] The recognition module is used to input the obtained frequency domain feature data into the pre-trained fully mechanized mining face voice control command recognition model for processing, generating text data. The training method for the fully mechanized mining face voice control command recognition model is as follows:
[0015] Step 1: Collect environmental noise and voice control commands from workers in the corresponding environments at the fully mechanized mining face, the central control center, and the ground dispatch center. At the same time, collect voice control commands from workers in quiet environments.
[0016] Step 2: Combine the noise data of the corresponding environment with the voice control command data collected in the absence of ambient noise, and label them together with the voice control commands recorded in the corresponding environment. Then, generate training, validation and test datasets through the enhancement module.
[0017] Step 3: Construct a voice control command recognition model for the fully mechanized mining face. The first part is a multi-layer convolutional neural network, the second part is a multi-layer recurrent neural network, and either a long short-term memory network or a gated recurrent unit network is selected. The third part is a fully connected network, and the fourth part is a decoder.
[0018] Step 4: Input the training set into the established fully mechanized mining face voice control command recognition model, observe the test results of the validation set, and continuously optimize the network complexity and improve the algorithm through repeated training. The model obtained when the test results of the test set meet the actual engineering requirements is used as the final fully mechanized mining face voice control command recognition model.
[0019] The display module displays the generated text data on the corresponding monitoring screen through a pre-designed interface, simultaneously broadcasts the generated text data aloud, and stores the audio data and generated text data in a database or log for later verification. The control module breaks down the generated text data according to the control object, control mode, and control action of the longwall mining face, generates corresponding control commands, and then sends them to the central control center. The central control center then operates the controlled object to perform actions based on the control commands.
[0020] The beneficial effects of this invention are as follows: The voice acquisition system of this invention has voice information input function and low power consumption function in monitoring mode; the voice processing system has voice data processing function and voice recognition function, and can prepare for recognition of voice commands under the special environmental noise of the fully mechanized mining face; the voice control system has the functions of displaying, broadcasting and storing, and generating and sending control commands. Users only need to issue a wake-up command and simultaneously issue a voice control command, and the system can display, broadcast and store the command in text form on the monitoring screen, while simultaneously realizing the remote control function of the fully mechanized mining face equipment. In the voice processing system of this invention, a series of voice noise reduction and voice enhancement processes are performed on the acquired voice data for the complex and special environment of the fully mechanized mining face, improving the accuracy of voice recognition under specific conditions. Furthermore, the integrated voice control system enhances the timeliness and convenience of remote control of the fully mechanized mining face equipment, reducing the occurrence of accidents such as misoperation. Attached Figure Description
[0021] Figure 1 This is a schematic diagram of the voice recognition control system for centralized control commands in fully mechanized mining faces according to the present invention.
[0022] Figure 2 A flowchart illustrating the enhancement process performed on the enhancement module of this invention;
[0023] Figure 3 This is a flowchart of the voice recognition control system for centralized control commands in fully mechanized mining faces according to the present invention. Detailed Implementation
[0024] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0025] like Figure 1 As shown, the present invention provides a voice recognition and control system for centralized control commands in a fully mechanized mining face, comprising a voice acquisition system, a voice processing system, and a voice control system; the voice acquisition system includes a monitoring module and an acquisition module, the voice processing system includes an enhancement module and a recognition module, and the voice control system includes a display module and a control module.
[0026] The monitoring module is used to maintain low-power operation when the system is in a sleep state and wait for a wake-up command. Before receiving a wake-up command, all other modules except the monitoring module are in a sleep state to reduce system power consumption. After receiving a wake-up command, the monitoring module wakes up other modules in the system and enters the working state.
[0027] The acquisition module is used to collect voice control commands issued by staff through handheld terminal devices, central control centers, or ground dispatch centers after receiving a wake-up command, and to convert the collected voice control commands into audio data that can be recognized by computers through analog-to-digital conversion.
[0028] The enhancement module is used to enhance the acquired audio data, obtain its frequency domain features, and realize speech enhancement, speech denoising, and feature extraction. For example... Figure 2 As shown, the specific method of enhancement processing is as follows:
[0029] Step 1: Use wavelet packet transform to convert the speech time-domain signal to the frequency domain, obtaining N-layer 2. N The wavelet packet decomposition algorithm for discrete signal x(t) is as follows: (The algorithm uses multiple frequency domain sub-signals from different frequency bands to enable better time-frequency localization analysis and extraction of local signal features.)
[0030]
[0031] In the formula, i is the number of decomposition layers, l is the translation factor, and v(2 i tl) is the basis function, ci,j is the j-th coefficient obtained from the i-th level wavelet packet decomposition, and t is the sampling period.
[0032] Step 2: Calculate the energy entropy of each sub-signal using its wavelet packet coefficients;
[0033] Calculate the energy value of each sub-signal and normalize it:
[0034]
[0035] In the formula, p i,j Let be the proportion of the energy of the j-th sub-signal in the i-th layer to the total energy in the i-th layer. It can also be seen that the energy of the wavelet packet decomposition node is the square of the wavelet packet coefficient.
[0036] By fusing wavelet packet multi-scale decomposition information entropy theory, the wavelet packet energy entropy of the signal can be calculated:
[0037] H i,j =-p i,j logp i,j (3)
[0038] In the formula, H i,jLet be the wavelet packet energy entropy of the j-th sub-signal in the i-th layer.
[0039] Step 3: Due to the diversity and singularity of noise in underground coal mines, the high and low frequency sub-signals of wavelet packet decomposition will have certain errors. Therefore, an adaptive threshold method is used to further divide the frequency band of each sub-signal of the wavelet packet to obtain the high frequency sub-signals.
[0040]
[0041] In the formula, σ is the estimated noise variance. MAD is the median of all high-frequency wavelet coefficients, and the energy entropy H is... i,j Greater than λ i The signal is the high-frequency sub-signal.
[0042] Step 4: The high correlation between noise signal data frames gives the background noise time-frequency matrix a low-rank structure. Compared to the background noise, the speech signal is often active only in a few frequency components, exhibiting a certain degree of sparsity. Therefore, low-rank and sparse matrix decomposition is applied to the high-frequency sub-signals to obtain the speech data and background noise data:
[0043] M = L + S (5)
[0044] M represents the noisy high-frequency sub-signal, S represents the effective speech information, and L represents the background noise information.
[0045] Step 5: Optimize the phase spectrum using the phase compensation function Λ, and recombine the amplitude spectrum of the noisy high-frequency sub-signal with the optimized phase spectrum to obtain the frequency domain features S of the enhanced speech signal. Λ :
[0046]
[0047] ∠M Λ =arg[M+Λ] (7)
[0048] S Λ =|M|exp(j∠M) Λ (8)
[0049] Here is the noise amplitude estimate, and λ is the compensation factor, taken as 3.74. Let ∠M be the decision function. Λ The optimized phase spectrum is given by |M|, which is the amplitude spectrum of the noisy high-frequency sub-signal. In formula (8), j is the imaginary unit.
[0050] The recognition module is used to input the obtained frequency domain data into the pre-trained fully mechanized mining face voice control command recognition model for processing, generating text data. The training method for the fully mechanized mining face voice control command recognition model is as follows:
[0051] Step 1: Collect environmental noise and voice control commands from workers in the corresponding environments at the fully mechanized mining face, the central control center, and the ground dispatch center. At the same time, collect voice control commands from workers in quiet environments.
[0052] Step 2: Combine the noise data of the corresponding environment with the voice control command data collected in the absence of ambient noise, and label them together with the voice control commands recorded in the corresponding environment. Then, generate training, validation and test datasets through the enhancement module.
[0053] Step 3: Construct a voice control command recognition model for the fully mechanized mining face. The first part is a multi-layer convolutional neural network; the second part is a multi-layer recurrent neural network. Considering the computing power, the multi-layer recurrent neural network is selected as a long short-term memory network or a gated recurrent unit network; the third part is a fully connected network; and the fourth part is a decoder.
[0054] The activation function of the multi-layer convolutional neural network is the GELU function, which not only preserves probabilistic properties but also retains dependence on the input, effectively avoiding the gradient vanishing problem. The formula for calculating the GELU function is:
[0055]
[0056] The formula for calculating the tanh function is as follows:
[0057]
[0058] Step 4: Input the training set into the established fully mechanized mining face voice control command recognition model, observe the test results of the validation set, and continuously optimize the network complexity and improve the algorithm through repeated training. The model obtained when the test results of the test set meet the actual engineering requirements is used as the final fully mechanized mining face voice control command recognition model.
[0059] The display module displays the generated text data on the corresponding monitoring screen through a pre-designed interface, simultaneously broadcasts the generated text data aloud, and stores the audio data and generated text data in a database or log for later verification. The control module breaks down the generated text data according to the control object, control mode, and control action of the longwall mining face, generates corresponding control commands, and then sends them to the central control center. The central control center then operates the controlled object to perform actions based on the control commands.
[0060] like Figure 3 As shown, the working principle of this invention is as follows: The voice recognition control system for centralized control commands at the fully mechanized mining face is normally in a dormant state, maintaining low power consumption and not processing external voice information. When a wake-up command is detected, the system collects and stores the audio data transmitted by the staff through the voice device, then processes and recognizes the audio data, displays the generated text data on the monitoring screen, broadcasts and stores it, and then further generates control commands and sends them to the centralized control center.
[0061] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.
Claims
1. A voice recognition control system for centralized control commands in a fully mechanized mining face, characterized in that, It includes a voice acquisition system, a voice processing system, and a voice control system; the voice acquisition system includes a monitoring module and a acquisition module, the voice processing system includes an enhancement module and a recognition module, and the voice control system includes a display module and a control module. The enhancement module is used to enhance the acquired audio data to obtain its frequency domain characteristics; The specific methods for enhancement processing are as follows: Step 1: Perform wavelet packet decomposition on the audio data; the wavelet packet decomposition algorithm for discrete signal x(t) is as follows: In the formula, i is the number of decomposition layers, l is the translation factor, and v(2 i tl) is a basis function, c i,j Let be the j-th coefficient obtained from the i-th level wavelet packet decomposition, and t be the sampling period; Step 2: Calculate the energy entropy corresponding to each sub-signal using its wavelet packet coefficients; calculate the energy value of each sub-signal and normalize it. In the formula, p i,j Let be the proportion of the energy of the j-th sub-signal in the i-th layer to the total energy in the i-th layer. It can also be seen that the energy of the wavelet packet decomposition node is the square of the wavelet packet coefficient. By fusing wavelet packet multi-scale decomposition information entropy theory, the wavelet packet energy entropy of the signal can be calculated: H i,j =-p i,j logp i,j (3) In the formula, H i,j Let the wavelet packet energy entropy be the i-th layer and j-th sub-signal. Step 3: Obtain the high-frequency sub-signal using an adaptive thresholding method: In the formula, σ is the estimated noise variance. MAD is the median of all high-frequency wavelet coefficients, and the energy entropy H is... i,j Greater than λ i The signal is the high-frequency sub-signal; Step 4: Perform low-rank and sparse matrix decomposition on the high-frequency sub-signals to obtain speech data and background noise data: M = L + S (5) M represents the noisy high-frequency sub-signal, S represents the effective speech information, and L represents the background noise information; Step 5: Optimize the phase spectrum using the phase compensation function Λ, and recombine the amplitude spectrum of the noisy high-frequency sub-signal with the optimized phase spectrum to obtain the frequency domain features S of the enhanced speech signal. Λ : ∠M Λ =arg[M+Λ] (7) S Λ =|M|exp(j∠M Λ ) (8) Here is the noise amplitude estimate, and λ is the compensation factor, taken as 3.
74. Let ∠M be the decision function. Λ The optimized phase spectrum is given by |M|, which is the amplitude spectrum of the noisy high-frequency sub-signal. In formula (8), j is the imaginary unit.
2. The voice recognition control system for centralized control commands in fully mechanized mining faces according to claim 1, characterized in that, The monitoring module is used to maintain low-power operation when the system is in a sleep state, wait for a wake-up command, and wake up other modules of the system after receiving the wake-up command. The acquisition module is used to collect voice control commands issued by staff through handheld terminal devices, central control centers, or ground dispatch centers after receiving a wake-up command, and to convert the collected voice control commands into audio data that can be recognized by computers through analog-to-digital conversion.
3. The voice recognition control system for centralized control commands in fully mechanized mining faces according to claim 1, characterized in that, The recognition module is used to input the obtained frequency domain feature data into the pre-trained fully mechanized mining face voice control command recognition model for processing, and generate text data.
4. The voice recognition control system for centralized control commands in fully mechanized mining faces according to claim 3, characterized in that, The training method for the fully mechanized mining face voice control command recognition model is as follows: Step 1: Collect environmental noise and voice control commands from workers in the corresponding environments at the fully mechanized mining face, the central control center, and the ground dispatch center. At the same time, collect voice control commands from workers in quiet environments. Step 2: Combine the noise data of the corresponding environment with the voice control command data collected in the absence of ambient noise, and label them together with the voice control commands recorded in the corresponding environment. Then, generate training, validation and test datasets through the enhancement module. Step 3: Construct a voice control command recognition model for the fully mechanized mining face. The first part is a multi-layer convolutional neural network; the second part is a multi-layer recurrent neural network, which can be either a long short-term memory network or a gated recurrent unit network; the third part is a fully connected network; and the fourth part is a decoder. Step 4: Input the training set into the established fully mechanized mining face voice control command recognition model, observe the test results of the validation set, and continuously optimize the network complexity and improve the algorithm through repeated training. The model obtained when the test results of the test set meet the actual engineering requirements is used as the final fully mechanized mining face voice control command recognition model.
5. The voice recognition control system for centralized control commands in fully mechanized mining faces according to claim 1, characterized in that, The display module is used to display the generated text data on the corresponding monitoring screen through a pre-designed display interface, while simultaneously broadcasting the generated text data via voice. The audio data and the generated text data are stored in the database and logs for later verification.
6. The voice recognition control system for centralized control commands in fully mechanized mining faces according to claim 1, characterized in that, The control module is used to break down the generated text data according to the control object, control mode and control action of the fully mechanized mining face, generate corresponding control instructions, and then send them to the central control center, which will then operate the controlled object to perform actions in a unified manner according to the control instructions.
Citation Information
Patent Citations
Single-channel monitor-free voice and noise separating method based on low-rank and sparse matrix decomposition
CN102915742A
Voice denoising processing method and device, computer device and storage medium
CN110797041A
Deep neural network speech recognition method based on speech enhancement in complex environment
CN111986661A
Intelligent control system and method for fully mechanized coal mining face
CN115016336A