A mobile patrol voice and semantic recognition system with anti-noise function
By designing a noise-resistant speech semantic recognition system in the mobile inspection system, and using audio detection and noise reduction modules to separate and clean voice data, the problems of weak noise resistance and high semantic recognition error rate in the mobile inspection system are solved, significantly improving the accuracy of speech recognition and the efficiency of human-computer interaction.
Patent Information
- Application Number
- CN202111315295.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-08
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-11-08
AI Technical Summary
The existing mobile inspection system has problems such as insufficient intelligence and poor human-computer interaction at the front end of the manual mobile inspection, especially in the case of weak noise resistance and high semantic recognition error rate, which affects the accuracy and efficiency of data acquisition.
A noise-resistant mobile patrol speech semantic recognition system is designed, including an audio acquisition module, an audio detection module, an audio noise reduction module and a semantic recognition module. The audio information is detected and marked through the audio detection module, voice data and noise data are separated, and transmitted to the audio noise reduction module for noise reduction processing, and clean voice data is generated for semantic recognition.
It effectively reduces noise interference in speech data, significantly improves the accuracy of speech recognition, improves the efficiency of human-computer interaction, and solves the problems of weak noise resistance and high semantic recognition error rate in the prior art.
Smart Images

Figure CN114171017B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of mobile inspection voice and semantic recognition, and particularly to a noise-resistant mobile inspection voice and semantic recognition system. Background Art
[0002] At present, the State Grid has initially realized the manual mobile operation of inspection services, with primary functions such as real-time viewing of line equipment ledgers and real-time reporting of defects and potential hazards. However, there are problems such as insufficient intelligence and poor human-computer interaction in on-site front-end operations, and most of the front-end data collection in manual mobile inspections relies on manual input. Although the front-end of manual mobile inspections supports voice interaction, the effect is not very good, mainly having problems such as weak noise resistance and high semantic recognition error rates. Summary of the Invention
[0003] The purpose of the present invention is to provide a noise-resistant mobile inspection voice and semantic recognition system, which reduces the noise interference in voice data and improves the accuracy of voice recognition.
[0004] To solve the above technical problems, the technical solution of the present invention is: a noise-resistant mobile inspection voice and semantic recognition system, including:
[0005] An audio acquisition module realizes the acquisition of audio information;
[0006] An audio detection module realizes the detection and annotation of audio information to obtain voice data and noise data;
[0007] An audio noise reduction module realizes the noise reduction of voice data;
[0008] A semantic recognition module realizes the semantic recognition of the noise-reduced voice data.
[0009] Preferably, the audio acquisition module remains always on, continuously acquires nearby audio information, and transmits the audio information acquired every 1 second to the audio detection module.
[0010] Preferably, the audio detection module detects whether there is voice input in the audio information. If no voice input is detected, the audio information is saved and marked as noise data; if voice input is detected, the audio information is saved and marked as voice data;
[0011] After detecting the voice data, all the stored noise data and the voice data are transmitted to the audio noise reduction module.
[0012] Preferably, the audio detection module saves the noise data with a length of 10 seconds, and each time the latest segment of the noise data is saved, it overwrites the earliest segment of the noise data.
[0013] Preferably, the audio noise reduction module performs noise reduction processing on the speech data to obtain generated speech data, and then transmits the generated speech data to the semantic recognition module. The specific steps include:
[0014] Step 1: Respectively take the noise data and the speech data as cut-off points at equal time intervals, and cut them into several audio data blocks;
[0015] Step 2: Convert each audio data block into a corresponding digital matrix and perform normalization processing to obtain noise digital data and speech digital data;
[0016] Step 3: Concatenate the noise digital data and the speech digital data, and then transmit the concatenated digital matrix into a trained speech generation model to obtain generated speech data.
[0017] Preferably, the speech generation model includes 2 downsampling structures and 1 upsampling structure, where each downsampling structure includes 3 down-convolution layers, and the upsampling structure includes 3 transposed convolution layers;
[0018] One of the downsampling structures downsamples the speech digital data, while the other downsampling structure downsamples the concatenated speech digital data and noise digital data; Concatenate the feature data obtained by the two downsamplings, and after the concatenated feature data is convolved twice and then upsampled, the result of each upsampling is concatenated with the feature of the downsampled speech digital data, and finally the generated speech data is obtained through multiple upsamplings of the upsampling structure.
[0019] Preferably, the semantic recognition module realizes semantic recognition of the generated speech data. The specific steps include:
[0020] Step 1: Preprocess the generated speech data, and extract the speech feature sequence that changes with time from the waveform of the generated speech data;
[0021] Step 2: Transmit the speech feature sequence into the established search space, and find the best word string through Viterbi search.
[0022] Preferably, the preprocessing process is to perform endpoint detection of the speech signal, speech framing and pre-emphasis processing on the generated speech data.
[0023] Preferably, the search space includes an acoustic model, a language model or a speech dictionary.
[0024] Preferably, the speech dictionary is a keyword database established for terms related to the power industry, including: power equipment names, commonly used words for power operations, basic power terms, and commonly used power unit names.
[0025] The beneficial effects of the present invention are as follows: compared with the prior art, the present invention uses an audio detection module to detect and label audio information. After obtaining voice data and noise data, it is sent to a voice generation model in the audio noise reduction module for noise reduction to obtain generated voice data. Based on the generated voice data, semantic recognition is performed through a semantic recognition module, effectively reducing noise interference in voice data at the front end of manual mobile inspection, greatly improving the accuracy of voice recognition, and improving the efficiency of human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a working flowchart of the anti-noise mobile inspection voice semantic recognition system according to an embodiment of the present invention;
[0027] Figure 2 It is a flowchart of noise reduction performed by the audio noise reduction module in this embodiment;
[0028] Figure 3 It is a network structure diagram of the voice generation model in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0030] Please refer to Figures 1 to 3 , the present invention is an anti-noise mobile inspection voice semantic recognition system, which includes: an audio acquisition module, an audio detection module, an audio noise reduction module, and a semantic recognition module;
[0031] The audio acquisition module realizes the acquisition of audio information;
[0032] The audio detection module realizes the detection and labeling of audio information to obtain voice data and noise data;
[0033] The audio noise reduction module realizes the noise reduction of voice data;
[0034] The semantic recognition module realizes the semantic recognition of the voice data after noise reduction.
[0035] Refer to Figure 1 , the audio acquisition module remains always on, continuously acquires nearby audio information, and transmits the audio information acquired every 1 second to the audio detection module.
[0036] The audio detection module detects whether there is voice input in the audio information. If no voice input is detected, the audio information is saved and marked as noise data; if voice input is detected, the audio information is saved and marked as voice data;
[0037] The audio detection module only saves the noise data with a length of 10 seconds. Each time the latest segment of noise data is saved, it will overwrite the earliest segment of noise data;
[0038] After the voice data is detected, all the stored noise data and voice data are transmitted to the audio noise reduction module.
[0039] See Figure 2 , the audio noise reduction module is used to perform noise reduction processing based on the noise data and voice data to obtain generated voice data, and then the generated voice data is transmitted to the semantic recognition module; the specific steps include:
[0040] Step 1: Respectively cut the noise data and voice data at equal time intervals into several audio data blocks; in this embodiment, the equal time interval is 20 ms;
[0041] Step 2: Convert each audio data block into a corresponding digital matrix and perform normalization processing to obtain noise digital data and voice digital data;
[0042] Step 3: Concatenate the noise digital data and the said voice digital data, and then transmit the concatenated digital matrix into a trained voice generation model to obtain generated voice data.
[0043] See Figure 3 , in step 3 of the above audio noise reduction module, the voice generation model includes 2 downsampling structures and 1 upsampling structure, where each downsampling structure includes 3 down-convolution layers, and the upsampling structure includes 3 transposed convolution layers; one of the downsampling structures downsamples the voice digital data, and at the same time the other downsampling structure downsamples the concatenated voice digital data and noise digital data; the feature data obtained by the two downsamplings are concatenated, and the obtained feature data is convolved twice and then upsampled. Each upsampling result is concatenated with the feature of the downsampled voice digital data, and finally the generated voice data is obtained through multiple upsamplings of the upsampling structure. The generated voice data has the same voice content as the voice data, but the generated voice data has lower noise or no noise.
[0044] The semantic recognition module realizes semantic recognition of the generated voice data, and the specific steps include:
[0045] Step 1: Preprocess the generated voice data to extract the voice feature sequence that changes with time from the waveform of the generated voice data; in this embodiment, the preprocessing process is to perform endpoint detection, voice framing and pre-emphasis processing on the generated voice data.
[0046] Step 2: Input the speech feature sequence into the established search space, and find the optimal word string through Viterbi search. In this embodiment, the search space includes an acoustic model, a language model, or a speech dictionary; among them, the speech dictionary is a keyword database established for various specialized terms related to the power industry, including: power equipment names, commonly used words for power operations, basic power terms, commonly used power unit names, etc.
[0047] Parts not involved in the present invention are the same as or implemented by the prior art.
[0048] The above uses specific specific examples to illustrate the implementation manners of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the above embodiments and the features in the embodiments can be combined with each other. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
Claims
1. An anti-noise mobile inspection voice semantic recognition system, characterized in that: It includes An audio acquisition module for acquiring audio information; An audio detection module for detecting and annotating audio information to obtain speech data and noise data; The audio detection module detects whether there is speech input in the audio information. If no speech input is detected, the audio information is saved and marked as noise data; If speech input is detected, the audio information is saved and marked as speech data; After detecting the speech data, all the stored noise data and the speech data are transmitted to the audio noise reduction module; The audio noise reduction module reduces the noise of the speech data; The audio noise reduction module performs noise reduction processing on the speech data to obtain generated speech data, and then transmits the generated speech data to the semantic recognition module; The specific steps include: Step 1: Respectively cut the noise data and the speech data at equal time intervals into several audio data blocks; Step 2: Convert each audio data block into a corresponding digital matrix and perform normalization processing to obtain noise digital data and speech digital data; Step 3: Concatenate the noise digital data and the speech digital data, and then transmit the concatenated digital matrix into a trained speech generation model to obtain generated speech data; The speech generation model includes 2 downsampling structures and 1 upsampling structure, where each downsampling structure includes 3 down-convolution layers, and the upsampling structure includes 3 transposed convolution layers; One of the downsampling structures downsamples the speech digital data, and at the same time the other downsampling structure downsamples the concatenated speech digital data and noise digital data; Concatenate the feature data obtained by the two downsamplings, and after the obtained feature data is convolved twice and then upsampled, the result of each upsampling is concatenated with the feature of the downsampled speech digital data, and finally the generated speech data is obtained through multiple upsamplings of the upsampling structure; The semantic recognition module realizes the semantic recognition of the noise-reduced speech data.
2. The anti-noise mobile inspection voice semantic recognition system according to claim 1, characterized in that: The audio acquisition module remains always on, continuously acquires nearby audio information, and transmits the audio information acquired every 1 second to the audio detection module.
3. The anti-noise mobile inspection voice semantic recognition system according to claim 1, characterized in that: The audio detection module saves the noise data with a length of 10 seconds, and each time the latest segment of the noise data is saved, it overwrites the earliest segment of the noise data.
4. The anti-noise mobile inspection voice semantic recognition system according to claim 1, characterized in that: The semantic recognition module realizes the semantic recognition of the generated speech data. The specific steps include: Step 1: Preprocess the generated speech data, and extract the speech feature sequence that changes with time from the waveform of the generated speech data; Step 2: Transmit the speech feature sequence into the built search space, and find the best word string through Viterbi search.
5. The anti-noise mobile inspection voice semantic recognition system according to claim 4, It is characterized in that: The preprocessing process performs endpoint detection of the speech signal, speech framing, and pre-emphasis processing on the generated speech data.
6. The anti-noise mobile inspection speech semantic recognition system according to claim 4, It is characterized in that: The search space includes an acoustic model, a language model, or a speech dictionary.
7. The anti-noise mobile inspection speech semantic recognition system according to claim 6, It is characterized in that: The speech dictionary is a keyword database established for terms related to the power industry, including: names of power equipment, common words for power operations, basic power terms, and names of common power units.
Citation Information
Patent Citations
Background noise-reduction optimization method during Sphinx speaking speed recognition process
CN107123419A